How to use Jev with Python
By Flavio Copes
How to use Jev with Python: install typesafe-sdk, ask Choice, Score and Noul questions, read the probabilities, and classify a whole CSV with asyncio.
To use Jev with Python, install the typesafe-sdk package, set the TYPESAFE_API_KEY environment variable, create a TypeSafeClient and call its system_one() method with a state and a dictionary of questions. The answers come back as typed objects you read with response.choices["topic"].choice, response.scores["frustration"].score and response.nouls["wants_reply"].noul.
Jev is TypeSafe AI’s decision model. Instead of generating text, it returns typed decisions: a yes/no probability, one option from a list you wrote, or a position on a scale you described, always with probabilities. My deep dive into Jev explains the model. This post is about the Python code, and if you work in JavaScript there’s a companion post on how to use Jev in Node.js.
We’ll build one small script that reads a CSV of customer feedback, asks Jev three questions about each row, and writes the answers to a new CSV. It starts as one synchronous call and ends as an async version for thousands of rows.
The quick answer
| What | How |
|---|---|
| Install | pip install typesafe-sdk or uv add typesafe-sdk |
| Requirements | Python 3.10 or newer, an API key from the TypeSafe console |
| API key | TYPESAFE_API_KEY, or TypeSafeClient(api_key=...) |
| Ask | client.system_one(state=..., questions={...}) |
| Question types | Choice, Score, Noul |
| Read answers | response.choices, response.scores, response.nouls |
| Pin the model | TypeSafeClient(model="jev-1.13.0") |
| Many rows | AsyncTypeSafeClient plus asyncio.gather and a semaphore |
What are we building?
Imagine a support inbox exported to feedback.csv:
id,message
1,"I was charged twice for the Pro plan this month. Please refund one of the charges."
2,"The CSV export has been timing out since Monday. We have a board meeting tomorrow and need those numbers."
3,"Would love a dark mode for the dashboard, my eyes thank you in advance."
4,"Just wanted to say the new onboarding is great. Set up our whole team in ten minutes."
5,"Third time I'm writing about this. The invoice PDF still shows our old address. Fix it or we cancel."
For each message we want to know what it’s about (a Choice between bug, billing, feature request, praise and other), how frustrated the customer is (a Score on three levels), and whether they expect a reply (a Noul, Jev’s yes/no question).
The script writes those answers to results.csv, with a needs_review column for the rows where Jev wasn’t sure.
How do I install the Jev Python SDK?
The official package is typesafe-sdk on PyPI, imported as typesafe_sdk. It needs Python 3.10 or newer, so check your version first:
python3 --version
With pip, create a project folder and a virtual environment, so the SDK and its dependencies stay inside this project:
mkdir feedback-triage
cd feedback-triage
python3 -m venv .venv
source .venv/bin/activate
pip install typesafe-sdk
In fish, activate it with source .venv/bin/activate.fish instead.
If you use uv, it creates the project and its virtual environment for you:
uv init feedback-triage
cd feedback-triage
uv add typesafe-sdk
Then run scripts with uv run triage.py, no activation needed. On macOS you can install uv with brew install uv. If your Python is older than 3.10, uv init --python 3.13 feedback-triage makes uv fetch a newer one for this project.
How do I set the API key?
You need a TypeSafe account to create a key in the TypeSafe console. As of late September 2026, TypeSafe has paused new signups because of demand, while existing accounts keep working, so check typesafe.ai for the current state. The steps are in how to get access to Jev and an API key.
The client reads it from the environment:
export TYPESAFE_API_KEY=your_key_here
In fish, use set -x TYPESAFE_API_KEY your_key_here.
With uv you can also keep the key in a .env file and run uv run --env-file .env triage.py, which loads it for you (the SDK doesn’t read .env files). Add .env to .gitignore, since the one uv init generates doesn’t list it.
If no key is found, TypeSafeClient() raises a TypeSafeError right away, before sending anything.
Every setting can also be passed to the constructor as a keyword argument, which wins over the environment variable:
| Argument | Environment variable | Default |
|---|---|---|
api_key | TYPESAFE_API_KEY | none, required |
model | TYPESAFE_DEFAULT_MODEL | jev-latest |
base_url | TYPESAFE_BASE_URL | https://api.typesafe.ai |
timeout | 10 seconds per HTTP operation | |
retry | RetryPolicy() defaults |
One more variable, TYPESAFE_LOG_LEVEL, turns on the SDK’s logging. info writes one line per request. Be careful with debug: it logs request and response bodies without redacting them, so every customer message ends up in your logs.
How do I make my first call?
Let’s classify a single message. system_one() takes the state, the data Jev looks at, and the questions, a dictionary where each key is an ID you pick and each value is a question.
Using the client in a with block closes its HTTP connections when the block ends. Without it, call client.close() when you’re done.
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
message = "I was charged twice for the Pro plan this month. Please refund one of the charges."
with TypeSafeClient() as client:
response = client.system_one(
state={"message": message},
questions={
"topic": Choice(
instructions="What is `message` mainly about?",
criteria={
"bug": "Something in the product is broken or behaves wrong",
"billing": "Charges, invoices, refunds, plans, or payment methods",
"feature_request": "Asks for something the product does not do yet",
"praise": "Says something positive and asks for nothing",
"other": None,
},
),
"frustration": Score(
instructions="How frustrated is the author of `message`?",
criteria=[
"Calm, just stating facts",
"Annoyed but polite",
"Angry, uses strong language, or threatens to cancel",
],
),
"wants_reply": Noul(
instructions="Does `message` ask us to reply or to take an action?",
),
},
)
print(response.choices["topic"].choice)
print(response.scores["frustration"].score)
print(response.nouls["wants_reply"].noul)
The backticks around message point each question at that field of the state, as the TypeSafe docs recommend. The question IDs are for your code and aren’t sent to the model, so the full question goes in instructions.
Choice takes a dictionary of options: each key is a label you can get back, each value describes when it applies, and None works for an other bucket. Score takes a list of levels from low to high, where the position of each level is its number, starting at 0. Noul only needs instructions, plus an optional criteria={"true": ..., "false": ...} when the line between yes and no is subtle.
Should I use question objects or plain dicts?
The SDK accepts both. Here are the same questions as plain dictionaries, in the shape the HTTP API expects:
from typesafe_sdk import Questions
QUESTIONS: Questions = {
"topic": {
"type": "choice",
"instructions": "What is `message` mainly about?",
"criteria": {
"bug": "Something in the product is broken or behaves wrong",
"billing": "Charges, invoices, refunds, plans, or payment methods",
"feature_request": "Asks for something the product does not do yet",
"praise": "Says something positive and asks for nothing",
"other": None,
},
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the author of `message`?",
"criteria": [
"Calm, just stating facts",
"Annoyed but polite",
"Angry, uses strong language, or threatens to cancel",
],
},
"wants_reply": {
"type": "noul",
"instructions": "Does `message` ask us to reply or to take an action?",
},
}
My advice is to use the objects. They’re pydantic models, so a typo like critera= raises a ValidationError the moment you build the question. For a dictionary, the SDK only checks that type is set and that Choice and Score have criteria, then sends it as written, unknown keys included.
Dictionaries make sense when the questions come from a JSON file or a database, and you can mix both styles in one request. Whichever you pick, annotate a questions constant with the SDK’s Questions type. Without it, mypy infers a dictionary of objects as dict[str, _Question] and rejects it when you pass it to system_one().
How do I read the answers?
system_one() returns a SystemOneResponse. The answers are grouped by type, so you look each one up by its question ID in choices, scores or nouls. The values in these comments are illustrative:
topic = response.choices["topic"]
print(topic.choice) # billing
print(topic.probabilities) # {'bug': 0.02, 'billing': 0.95, 'feature_request': 0.0, 'praise': 0.0, 'other': 0.03}
print(topic.confidence) # 0.91
frustration = response.scores["frustration"]
print(frustration.score) # 1.1
print(frustration.probabilities) # {0: 0.05, 1: 0.8, 2: 0.15}
print(frustration.legend[round(frustration.score)]) # Annoyed but polite
print(response.nouls["wants_reply"].noul) # 0.97
choice is the label with the highest probability, and probabilities has every label. score is the probability-weighted average of the level numbers, so it can land between levels. confidence goes from 0 to 1: high when the probability sits on one option, low when it’s spread out. A Noul has no confidence, because noul itself is the probability of yes.
Notice that Score probabilities and legend use integer keys in Python, even though the JSON has strings. That’s why frustration.legend[round(frustration.score)] gives you the description of the nearest level.
response.answers also holds every answer in one dictionary, but its values are a union of the three answer types, so a type checker wants an isinstance check before .choice. The grouped dictionaries skip that step.
The response carries some metadata too:
print(response.model) # jev-1.13.0
print(response.usage.input_tokens) # 118
print(response.usage.output_tokens) # 9
print(response.request_id)
model is the versioned ID that answered, even when you asked for jev-latest. usage.input_tokens is what you pay for, $0.042 per million as of September 2026, while output tokens are free. I explain how to estimate a bill in How much does Jev cost?. Both usage fields are typed int | None, so add or 0 before doing math with them. request_id identifies the request, useful when you report a problem to TypeSafe.
How do I pin the Jev model version?
The SDK default is jev-latest, an alias that as of September 2026 points to jev-1.13.0 (so does jev-preview). When TypeSafe ships a new release the alias moves, and your answers and confidence values can change without any change in your code. Once you’ve tuned thresholds against a version, pin it:
from typesafe_sdk import TypeSafeClient
client = TypeSafeClient(model="jev-1.13.0")
You can also pass model="jev-1.13.0" to a single system_one() call, or set TYPESAFE_DEFAULT_MODEL. Either way, log response.model next to each result.
To see which names your account can use, list the models:
from typesafe_sdk import TypeSafeClient
with TypeSafeClient() as client:
for model in client.models.list().models:
print(model.name, model.release_date, model.description)
The models page says the list currently contains only the aliases, and versioned IDs like jev-1.13.0 work anyway.
How do I process the whole CSV?
Now let’s turn the first call into a script, triage.py. The questions move into a QUESTIONS constant, and a to_row() function flattens each response into one CSV row:
import csv
import sys
from typesafe_sdk import Choice, Noul, Questions, Score, SystemOneResponse, TypeSafeClient
MODEL = "jev-1.13.0"
REVIEW_BELOW = 0.6
FIELDS = ["id", "topic", "topic_confidence", "frustration", "wants_reply", "needs_review", "model"]
QUESTIONS: Questions = {
"topic": Choice(
instructions="What is `message` mainly about?",
criteria={
"bug": "Something in the product is broken or behaves wrong",
"billing": "Charges, invoices, refunds, plans, or payment methods",
"feature_request": "Asks for something the product does not do yet",
"praise": "Says something positive and asks for nothing",
"other": None,
},
),
"frustration": Score(
instructions="How frustrated is the author of `message`?",
criteria=[
"Calm, just stating facts",
"Annoyed but polite",
"Angry, uses strong language, or threatens to cancel",
],
),
"wants_reply": Noul(
instructions="Does `message` ask us to reply or to take an action?",
),
}
def to_row(feedback: dict[str, str], response: SystemOneResponse) -> dict[str, str]:
topic = response.choices["topic"]
frustration = response.scores["frustration"]
wants_reply = response.nouls["wants_reply"]
unsure = topic.confidence < REVIEW_BELOW or frustration.confidence < REVIEW_BELOW
return {
"id": feedback["id"],
"topic": topic.choice,
"topic_confidence": f"{topic.confidence:.2f}",
"frustration": f"{frustration.score:.2f}",
"wants_reply": f"{wants_reply.noul:.2f}",
"needs_review": "yes" if unsure else "no",
"model": response.model,
}
def main(input_path: str, output_path: str) -> None:
with open(input_path, newline="", encoding="utf-8") as file:
rows = list(csv.DictReader(file))
with TypeSafeClient(model=MODEL) as client:
results = []
for row in rows:
response = client.system_one(state={"message": row["message"]}, questions=QUESTIONS)
results.append(to_row(row, response))
with open(output_path, "w", newline="", encoding="utf-8") as file:
writer = csv.DictWriter(file, fieldnames=FIELDS)
writer.writeheader()
writer.writerows(results)
if __name__ == "__main__":
main(sys.argv[1], sys.argv[2])
Run it with the input and output paths:
python triage.py feedback.csv results.csv
The state holds only the message, because that’s all the questions need. The 0.6 in REVIEW_BELOW is a starting point: run the script on rows you’ve labeled by hand and move it until the flagged rows are the ones that deserve a human.
This version makes one call at a time, and TypeSafe quotes 70 to 500 milliseconds per call (plus your network distance from its US West Coast servers). That’s fine for a few hundred rows. For more we want parallel calls, but first let’s handle failures.
What happens when a call fails?
The SDK retries some failures on its own and raises an exception when retries run out or can’t help. Everything inherits from TypeSafeError:
| Exception | When |
|---|---|
TypeSafeError | Base class, also raised before sending: missing key, empty questions, a Score with no levels |
TypeSafeAPIError | Any unsuccessful HTTP response, with .status, .body, .request_id |
TypeSafeAuthenticationError | 401, missing or wrong key |
TypeSafeUnprocessableEntityError | 422, the request failed validation, .body says which field |
TypeSafeRateLimitError | 429, with .retry_after_ms when the server sent it |
TypeSafeBadRequestError, TypeSafePermissionDeniedError, TypeSafeNotFoundError | 400, 403, 404 |
TypeSafeInternalServerError | Any 5xx, including 529 when TypeSafe is overloaded |
TypeSafeAPIResponseValidationError | A successful response with a malformed body, .field_path points at it |
TypeSafeAPIConnectionError | No HTTP response at all |
TypeSafeAPITimeoutError | The request exceeded its timeout |
The HTTP errors are subclasses of TypeSafeAPIError, and the timeout error is a subclass of the connection error, so catch the specific ones first:
import sys
from typesafe_sdk import (
TypeSafeAPIConnectionError,
TypeSafeAPIError,
TypeSafeAuthenticationError,
TypeSafeRateLimitError,
TypeSafeUnprocessableEntityError,
)
try:
response = client.system_one(state={"message": message}, questions=QUESTIONS)
except TypeSafeAuthenticationError:
sys.exit("Check your TYPESAFE_API_KEY")
except TypeSafeUnprocessableEntityError as error:
print("Invalid request:", error.body)
except TypeSafeRateLimitError as error:
print("Still rate limited after retries, wait", error.retry_after_ms, "ms")
except TypeSafeAPIError as error:
print("API error", error.status, error.request_id)
except TypeSafeAPIConnectionError:
print("Could not reach the TypeSafe API")
In the final script, an authentication error stops everything, since every row would fail the same way. Any other error only marks its own row.
How do I configure retries?
By default the SDK retries up to 2 times after the first attempt, on 408, 429 and every 5xx status, plus connection errors and timeouts. The wait starts at 0.5 seconds and doubles up to 5, with some jitter, unless the server sends retry-after or retry-after-ms, which the SDK follows. The whole call, retries included, gets 30 seconds.
A batch job can afford more patience, so pass a RetryPolicy:
from typesafe_sdk import RetryPolicy, TypeSafeClient
retry = RetryPolicy(max_retries=5, backoff_max=20.0, timeout=120.0)
client = TypeSafeClient(model="jev-1.13.0", retry=retry)
timeout here is the total budget in seconds, not the per-request timeout of the client. You can also pass retry= to a single call, and RetryPolicy(max_retries=0) turns retries off.
Be careful with http_statuses. It replaces the default set, and the example in the SDK reference, {429, 500, 502, 503, 504}, leaves out 529, the status TypeSafe returns when it’s overloaded.
How do I classify thousands of rows with the async client?
AsyncTypeSafeClient takes the same arguments and has the same methods, you only await them. With asyncio.gather we can run many calls at the same time, but not all at once: as of September 2026 the limits are 1,200 requests per minute and 250,000 tokens per second, adjusted dynamically during early access, and going over returns a 429.
A semaphore caps how many requests are in flight, but not the rate. With calls of about 100 milliseconds, 10 slots could start around 100 requests a second, five times what 1,200 a minute allows. So each row also waits index / STARTS_PER_SECOND seconds before it starts. At 15 starts per second that’s at most 900 requests a minute, leaving room for retries.
Here’s the function that classifies one row:
async def classify(
client: AsyncTypeSafeClient,
semaphore: asyncio.Semaphore,
index: int,
feedback: dict[str, str],
) -> dict[str, str]:
await asyncio.sleep(index / STARTS_PER_SECOND)
async with semaphore:
try:
response = await client.system_one(
state={"message": feedback["message"]},
questions=QUESTIONS,
)
except TypeSafeAuthenticationError:
raise
except TypeSafeError as error:
return failed_row(feedback, error)
return to_row(feedback, response)
And main() creates the client with async with, which closes it at the end (await client.aclose() does the same by hand), then gathers one task per row:
async def main(input_path: str, output_path: str) -> None:
with open(input_path, newline="", encoding="utf-8") as file:
rows = list(csv.DictReader(file))
semaphore = asyncio.Semaphore(MAX_IN_FLIGHT)
retry = RetryPolicy(max_retries=5, backoff_max=20.0, timeout=120.0)
async with AsyncTypeSafeClient(model=MODEL, retry=retry) as client:
results = await asyncio.gather(
*(classify(client, semaphore, i, row) for i, row in enumerate(rows))
)
gather returns the results in input order, so the output CSV lines up with the input. If the authentication error is raised, gather passes it up and asyncio.run() cancels the tasks still waiting.
At 15 starts per second, 10,000 rows take about 11 minutes. On a plan with higher limits, raise STARTS_PER_SECOND and MAX_IN_FLIGHT together.
Can I get typed attributes instead of dictionary lookups?
Yes. Subclass SystemOneResponse with one field per question ID and pass it as response_model:
from typesafe_sdk import ChoiceAnswer, NoulAnswer, ScoreAnswer, SystemOneResponse
class FeedbackResponse(SystemOneResponse):
topic: ChoiceAnswer
frustration: ScoreAnswer
wants_reply: NoulAnswer
response = client.system_one(
state={"message": message},
questions=QUESTIONS,
response_model=FeedbackResponse,
)
print(response.topic.choice, response.frustration.score, response.wants_reply.noul)
Now response.topic.choice autocompletes, and a missing or wrongly shaped answer raises TypeSafeAPIResponseValidationError right away instead of a KeyError later.
The complete script
Here’s the final triage.py, with the async client, the rate limiting, per-row errors, and a token count at the end:
import asyncio
import csv
import sys
from typesafe_sdk import (
AsyncTypeSafeClient,
Choice,
Noul,
Questions,
RetryPolicy,
Score,
SystemOneResponse,
TypeSafeAuthenticationError,
TypeSafeError,
)
MODEL = "jev-1.13.0"
MAX_IN_FLIGHT = 10
STARTS_PER_SECOND = 15
REVIEW_BELOW = 0.6
PRICE_PER_MILLION_INPUT_TOKENS = 0.042
FIELDS = [
"id",
"topic",
"topic_confidence",
"frustration",
"wants_reply",
"needs_review",
"model",
"input_tokens",
"error",
]
QUESTIONS: Questions = {
"topic": Choice(
instructions="What is `message` mainly about?",
criteria={
"bug": "Something in the product is broken or behaves wrong",
"billing": "Charges, invoices, refunds, plans, or payment methods",
"feature_request": "Asks for something the product does not do yet",
"praise": "Says something positive and asks for nothing",
"other": None,
},
),
"frustration": Score(
instructions="How frustrated is the author of `message`?",
criteria=[
"Calm, just stating facts",
"Annoyed but polite",
"Angry, uses strong language, or threatens to cancel",
],
),
"wants_reply": Noul(
instructions="Does `message` ask us to reply or to take an action?",
),
}
def to_row(feedback: dict[str, str], response: SystemOneResponse) -> dict[str, str]:
topic = response.choices["topic"]
frustration = response.scores["frustration"]
wants_reply = response.nouls["wants_reply"]
unsure = topic.confidence < REVIEW_BELOW or frustration.confidence < REVIEW_BELOW
return {
"id": feedback["id"],
"topic": topic.choice,
"topic_confidence": f"{topic.confidence:.2f}",
"frustration": f"{frustration.score:.2f}",
"wants_reply": f"{wants_reply.noul:.2f}",
"needs_review": "yes" if unsure else "no",
"model": response.model,
"input_tokens": str(response.usage.input_tokens or 0),
"error": "",
}
def failed_row(feedback: dict[str, str], error: TypeSafeError) -> dict[str, str]:
row = dict.fromkeys(FIELDS, "")
row.update(id=feedback["id"], needs_review="yes", error=type(error).__name__)
return row
async def classify(
client: AsyncTypeSafeClient,
semaphore: asyncio.Semaphore,
index: int,
feedback: dict[str, str],
) -> dict[str, str]:
await asyncio.sleep(index / STARTS_PER_SECOND)
async with semaphore:
try:
response = await client.system_one(
state={"message": feedback["message"]},
questions=QUESTIONS,
)
except TypeSafeAuthenticationError:
raise
except TypeSafeError as error:
return failed_row(feedback, error)
return to_row(feedback, response)
async def main(input_path: str, output_path: str) -> None:
with open(input_path, newline="", encoding="utf-8") as file:
rows = list(csv.DictReader(file))
semaphore = asyncio.Semaphore(MAX_IN_FLIGHT)
retry = RetryPolicy(max_retries=5, backoff_max=20.0, timeout=120.0)
async with AsyncTypeSafeClient(model=MODEL, retry=retry) as client:
results = await asyncio.gather(
*(classify(client, semaphore, i, row) for i, row in enumerate(rows))
)
with open(output_path, "w", newline="", encoding="utf-8") as file:
writer = csv.DictWriter(file, fieldnames=FIELDS)
writer.writeheader()
writer.writerows(results)
tokens = sum(int(row["input_tokens"] or 0) for row in results)
review = sum(row["needs_review"] == "yes" for row in results)
cost = tokens / 1_000_000 * PRICE_PER_MILLION_INPUT_TOKENS
print(f"{len(results)} rows, {review} to review, {tokens} input tokens, ${cost:.6f}")
if __name__ == "__main__":
asyncio.run(main(sys.argv[1], sys.argv[2]))
Run it as before, with python triage.py feedback.csv results.csv or uv run triage.py feedback.csv results.csv. It prints the row count, the rows to review, and the input tokens with their cost at the September 2026 price. Failed rows keep their id and carry the exception name in error, so you can rerun just those.
Before trusting the labels, open results.csv and check 50 or so rows by hand. When a label is wrong, fix the criteria for that option before you touch the threshold. The Python SDK reference is at docs.typesafe.ai/sdk/python, and the source is on GitHub.
Want me to talk about your product? You can sponsor this site.
Related posts about ai: