walid@portfolio:~/lab/jev-decision-model$
cd../lab
04ideaSep 2026

Seven jobs for a model that only decides

TypeSafe’s Jev never writes a sentence. It picks, scores and answers yes or no about text, for $0.042 per million tokens. Seven ways to use it, each with the exact question, the setup in ManyChat, Clay, Zapier, LangChain or LiteLLM, and what it costs

Jev is a model with one job: deciding things about text. Give it some text and a question and it picks an option, places the text on a scale or answers yes or no, always with probabilities and never with a sentence. That is why it is fast, and why it costs $0.042 per million input tokens with output free. Seven ways to put it to work follow, from finding the buyers in your Instagram comments to stopping an agent before it emails 4,212 contacts. Each comes with the exact question to ask, the setup and the cost, and the starter kit’s code is all on the page. Checked against the sources it cites, most of the numbers hold. Three don’t: the cascade result it headlines is within noise, invoices aren’t Jev’s weakest area, and Clay can’t hold 100,000 leads in one table.

JevClassificationLead scoringGuardrailsModel routingTypeSafe docs ↗Known weak spots ↗Jev on OpenRouter ↗Zapier app ↗LangChain integration ↗LiteLLM auto router ↗Langfuse evaluator ↗Every’s test ↗
One decision model, seven jobs around it. The 200× and 400× round TypeSafe’s own best case, 193.6× faster and 444.6× cheaper; Every’s independent test measured about 25× and 580×.
i
One job: deciding things about text

You give Jev some text and a question, and it answers in one of three shapes. Every answer comes back with probabilities, so you always know how sure it is. It never writes a sentence, and that is exactly why it is so fast and so cheap. The rule of thumb: if the output you want is a label, a score or a yes or no, use Jev. If you need words written, use a writing model, then use Jev to check the words.

The three shapes, and Every’s timing: 8.8 seconds for Claude Fable 5.1 against 0.35 for Jev, judging the same piece of writing.

The three shapes of answer

ShapeWhat comes backExample question
Pick oneWhat comes backOne option from a list you write, up to 255 options, with a probability for each.Example question“Is this reply interested, not now, or unsubscribe?”
ScoreWhat comes backA place on a scale of 2 to 10 levels that you describe.Example question“How good a fit is this lead?”
Yes or noWhat comes backOne number, the probability of yes. TypeSafe calls these “noul” questions.Example question“Does this draft make a claim we can’t back up?”
Checked, not copied

What held up, on 29 September 2026

Checked against TypeSafe’s docs, evals and launch post, and against the Every, OpenRouter, Langfuse, LiteLLM, LangChain, Vercel, Zapier, ManyChat and Clay pages the guide relies on.

ClaimWhat the primary source says
The priceCorrect: $0.042 per million input tokens, and output is free. TypeSafe’s pages mention no free tier; Zapier’s write-up says TypeSafe gives $5 of free credits a month.
AvailabilityJev launched on 15 September 2026, in what TypeSafe calls early access. Its launch post admits: “We can’t prove it isn’t subsidized.”
“25 times faster, 580 times cheaper”Correct, but a small test. Every used 12 synthetic passages with 7 planted mistakes. Jev took a median 0.35 seconds against 8.83 for Fable 5.1 at high effort, was an estimated 580 times cheaper, and caught 6 of the 7 where Fable caught all 7.
“193.6× faster, 444.6× cheaper”Quoted correctly. They come from TypeSafe’s own workflows, and TypeSafe itself calls them “the higher end of real world gains”.
“Weakest on invoices”Not quite. On TypeSafe’s own evals Jev scores 61.8% on invoices against 79.1% for OpenAI’s Sol, but security incidents is lower still at 61.7%. Invoices is where it trails the best model by the most, 17.3 points.
Langfuse’s 91.5%Correct, but it measures agreement with Fable 5.1, not accuracy. In the same 6,003-check study, DeepSeek V4.1 Flash agreed 93.5% of the time, at $260 per million against Jev’s $160.
The cascadeMisleading. On OpenRouter’s 50-question test the cascade made no mistakes for $0.012 and the big model alone made 2 for $0.175. But the cheap model alone also made none, for $0.004, and OpenRouter calls a two-answer gap between runs noise.
LiteLLM’s router testCorrect: 127 ms against 688 ms for Claude Haiku 4.5, a 96% cheaper classifier, and 95% against 74% agreement with the expected tiers. LiteLLM wrote the 80 prompts and their tiers itself, and graded the routing, not the answers.
“Same body” on OpenRouterHolds. Its Decisions endpoint takes the same fields and returns the same answers, plus the cost. OpenRouter also runs an exact copy of TypeSafe’s API at /api/v1/systemone for the TypeSafe SDKs.
The Zapier appReal, and new: Zapier’s own blog said on 24 September there was no native integration yet. It has one action, Ask Questions, which also takes a scale question the guide skips. The only choice for an unsure answer is Stop the Zap.
The agent guardStricter than it sounds. LangChain’s middleware blocks any call it rates 0.5 or more likely to be risky, a threshold you can’t change, and blocks too if the check fails. The package is an early alpha, 0.0.1a3.
Clay at 100,000 leadsWon’t fit in one table. Clay caps every table at 50,000 rows on every plan, and an import past the cap stops without an error.
Your dataTypeSafe doesn’t train on requests or responses, but zero data retention is for enterprise customers only. The LangChain guard sends up to 30 recent messages with each check, and LiteLLM sends earlier user turns by default.
Limits1,200 requests a minute and 64,000 tokens a request, of which 32,000 for the text plus the longest question. TypeSafe says the limits can change without notice, and English is where Jev is most accurate.
The numbers

What it costs, and how fast it is

Jev costs $0.042 per million input tokens, and output is free. In an independent test by Every, it judged a piece of writing in a median of 0.35 seconds against 8.83 seconds for Claude Fable 5.1, roughly 25 times faster, at an estimated 580 times lower cost. In the same test it caught 6 of 7 planted mistakes, and Fable caught all 7.

TypeSafe’s own figures go higher, up to 193.6 times faster and 444.6 times cheaper on its workflows, and it describes those as the high end of what you’ll see. It quotes 70 to 500 milliseconds end to end, served from the US West Coast.

Start here

Get access in five minutes

Pick the route that fits how you work.

RouteBest forWhat you need
Zapier, the “TypeSafe Jev” appBest forNo code at allWhat you needA Zapier account. Use the Ask Questions action.
OpenRouterBest forA quick start with no waitlistWhat you needAn OpenRouter key. Model typesafe/jev-1.13.
TypeSafe directBest forApps in productionWhat you needA key from console.typesafe.ai/keys. Model jev-1.13.0.
Vercel AI GatewayBest forApps already on VercelWhat you needModel typesafe-ai/jev, through the AI SDK 7 experimental evaluate API. It isn’t available on the Gateway’s OpenAI-compatible endpoints.

Your first call, straight to TypeSafe

Through OpenRouter the body is the same: swap the URL for https://openrouter.ai/api/alpha/decisions, use your OpenRouter key, and set the model to typesafe/jev-1.13.

curl.sh23 lines
curl https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-1.13.0",
    "state": "I was charged twice this month. Please fix it today.",
    "questions": {
      "topic": {
        "type": "choice",
        "instructions": "What is this message about?",
        "criteria": {
          "billing": "Payments, invoices or refunds",
          "technical": "Bugs or something not working",
          "sales": "Pricing or buying questions",
          "other": "None of the above"
        }
      },
      "is_urgent": {
        "type": "noul",
        "instructions": "The message asks for something to happen today or right away."
      }
    }
  }'

What comes back

“choice” is the winning option, and “confidence” says how clear the win was. A yes-or-no question returns one number, the probability of yes. Your numbers will differ.

response.json12 lines
{
  "model": "jev-1.13.0",
  "answers": {
    "topic": {
      "type": "choice",
      "choice": "billing",
      "confidence": 0.9,
      "probabilities": { "billing": 0.95, "technical": 0.03, "sales": 0.01, "other": 0.01 }
    },
    "is_urgent": { "type": "noul", "noul": 0.97 }
  }
}

The same thing through the SDK

pip install typesafe-sdk (Python 3.10 or later), or npm install @typesafe-ai/sdk (Node 20 or later), where the method is client.systemOne(). Both read TYPESAFE_API_KEY, default to jev-latest and retry on errors.

first_call.py16 lines
from typesafe_sdk import Choice, Noul, TypeSafeClient

with TypeSafeClient() as client:
    response = client.system_one(
        state={"document": "I was charged twice. Please fix this ASAP."},
        questions={
            "billing": Noul(instructions="Is this ticket about billing?"),
            "tone": Choice(
                instructions="What is the customer's tone?",
                criteria={"calm": None, "frustrated": None, "angry": None},
            ),
        },
    )

print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)

Check your key works

From the starter kit. It asks the same question through TypeSafe or OpenRouter, whichever key you have set.

first-call.sh22 lines
#!/usr/bin/env bash
# Your first Jev call. Set ONE of these first:
#   export TYPESAFE_API_KEY=...     (key from https://console.typesafe.ai/keys)
#   export OPENROUTER_API_KEY=...   (no TypeSafe account needed)
set -e
BODY='{
  "state": "I was charged twice this month. Please fix it today.",
  "questions": {
    "topic": { "type": "choice", "instructions": "What is this message about?",
      "criteria": { "billing": "Payments, invoices or refunds", "technical": "Bugs or something not working", "sales": "Pricing or buying questions", "other": "None of the above" } },
    "is_urgent": { "type": "noul", "instructions": "The message asks for something to happen today or right away." }
  }
}'
if [ -n "$TYPESAFE_API_KEY" ]; then
  echo "$BODY" | python3 -c 'import json,sys; b=json.load(sys.stdin); b["model"]="jev-1.13.0"; print(json.dumps(b))' |
  curl -s https://api.typesafe.ai/v1/systemone -H "Authorization: Bearer $TYPESAFE_API_KEY" -H "Content-Type: application/json" -d @- | python3 -m json.tool
elif [ -n "$OPENROUTER_API_KEY" ]; then
  echo "$BODY" | python3 -c 'import json,sys; b=json.load(sys.stdin); b["model"]="typesafe/jev-1.13"; print(json.dumps(b))' |
  curl -s https://openrouter.ai/api/alpha/decisions -H "Authorization: Bearer $OPENROUTER_API_KEY" -H "Content-Type: application/json" -d @- | python3 -m json.tool
else
  echo "Set TYPESAFE_API_KEY or OPENROUTER_API_KEY first."; exit 1
fi
Before any use case

How to write a question Jev gets right

Most bad results come from the question, so these rules matter more than anything else here.

✓Name the parts of your text. Send it as named fields, then point at them in backticks: “Is `comment` asking about `offer`?”✓Describe every option in a sentence. Jev reads the descriptions, so “interested: wants to learn more, asks a buying question, or says yes to a call” beats a bare “interested”✓Always add an escape option. Jev always picks an answer, so give it “other” or “unclear” for when nothing fits✓Write score levels as situations. List them lowest first and describe what each level looks like in the real world; numbers inside the level text don’t help✓Phrase a yes-or-no question so that yes is the thing you are hunting for. “The draft makes a claim that isn’t in the notes” is easier to act on than its opposite✓Keep numbers out of Jev. TypeSafe says plainly that it can’t count and isn’t a calculator; headcount filters, word counts and maths belong in a formula or code✓Set thresholds on your side, because the API has no “when unsure” setting. A sensible start: act above 0.8, send 0.5 to 0.8 to a person, drop the rest. For yes or no, treat above 0.8 as yes, below 0.2 as no, and look at the middle yourself✓Pin the version once it works. jev-latest moves when TypeSafe ships an update, so pin jev-1.13.0 once your thresholds are tuned
Use 1 of 7 · Deals

Find the buyers hiding in your comments and DMs

Every comment and DM is read the second it lands. 10,000 comments cost under 20 cents.

Every comment and DM that comes into ManyChat goes to Jev with one question: how close is this person to buying? Buyers get tagged and followed up in seconds, and fans get your normal reply. You stop scrolling through “🔥🔥🔥” to find the one person asking for a price. If you only set up one of the seven, make it this one; it is the fastest win if you sell anything on Instagram.

You need ManyChat on a paid plan, because the External Request step is a paid feature and only calls HTTPS links, plus a Jev key, and optionally Zapier and HubSpot.

Setup in ManyChat

1

Create three custom user fields

jev_intent (text), jev_intent_conf (number) and jev_asks_price (number).

2

Add the request

Open your comment or DM automation, add an Action block, then Automation › Make External Request › Add your request.

3

Point it at Jev

Request type POST, URL https://api.typesafe.ai/v1/systemone. Headers: Authorization set to Bearer YOUR_KEY, and Content-Type set to application/json.

4

Paste the body

Use the body below and change the offer line to what you sell.

5

Encode the message

When you insert Last text input, remove the quote marks around it and tick Encode to JSON. That stops a quote mark inside someone’s message from breaking the request.

6

Map the response

On the Response mapping tab, save $.answers.buying_intent.choice to jev_intent, $.answers.buying_intent.confidence to jev_intent_conf and $.answers.asks_price.noul to jev_asks_price.

7

Test it

Pick a test contact and hit Test Request on the Response tab. You want a 200 back. ManyChat doesn’t save mapped values during a test, so send one real message to check the fields fill.

8

Route on the answer

Add a Condition block. If jev_intent is ready_to_buy and jev_intent_conf is above 0.6, add the tag hot-lead and send your booking DM. If it’s considering, tag warm-lead and send one qualifying question. Everything else carries on through your normal flow.

9

Optional: your CRM

In Zapier, use the ManyChat trigger New Tagged User (tag hot-lead) and create the contact in HubSpot. Instagram contacts usually come without an email, so either ask for one in the DM first, or create the contact and keep the Instagram username in a custom property.

The request body

Replace the offer line. {{Last text input}} stays unquoted, because Encode to JSON adds the quotes.

01-buyer-intent-manychat-body.txt25 lines
{
  "model": "jev-1.13.0",
  "state": {
    "message": {{Last text input}},
    "instagram_username": "{{Instagram username}}",
    "offer": "REPLACE: one line on what you sell, e.g. Done-for-you AI automation for B2B service companies"
  },
  "questions": {
    "buying_intent": {
      "type": "choice",
      "instructions": "How close is the sender of `message` to buying `offer`?",
      "criteria": {
        "ready_to_buy": "Asks how to buy, the price, how to start, or to book a call",
        "considering": "Asks whether it would work for their business, about results, or how it works",
        "wants_freebie": "Only wants the free resource, template, or guide from the post",
        "existing_customer": "Already a customer asking for help",
        "just_engaging": "Praise, emoji, tagging a friend, or a general reaction",
        "spam": "Promotion, bot text, or unrelated selling",
        "unclear": "Too short or vague to tell"
      }
    },
    "asks_price": { "type": "noul", "instructions": "`message` asks about price, cost, or budget." },
    "has_business": { "type": "noul", "instructions": "`message` says the sender runs or works at a business." }
  }
}

Check this in your account: Last text input is the right field for DMs. For comment automations, open the variable picker and send a real test comment to confirm which field holds the comment text.

What it costs: about 400 input tokens a message, counting the question. 10,000 comments come to around 4 million tokens, about 17 cents.

Use 2 of 7 · Sales

Score 100,000 leads for under $3

Your team only calls the top of the list. The staff counts are for show: the guide itself says to filter headcount in Clay, never inside the question.

Every lead in your list gets a fit score from 1 to 5 plus a couple of yes-or-no flags, and the best ones land in HubSpot. Your team only calls the top of the list.

You need Apollo or any other lead list, Clay, a Jev key and HubSpot.

Setup in Clay

1

Export the leads

In Apollo, run your search, select the people and Export to CSV.

2

Import into Clay

Tools › Import › Import from CSV, and map the columns. Clay caps every table at 50,000 rows on every plan, and an import past the cap stops without an error, so split 100,000 leads across two tables.

3

Add an HTTP API column

Add enrichment, search HTTP API and open the Configure tab.

4

Store the key as a header account

Under Select header account, add an account with the header Authorization and the value Bearer YOUR_KEY, and name it TypeSafe Jev. Use a header account rather than the plain headers field, because Clay shows plain headers to anyone with access to the table.

5

Set the request

Method POST, endpoint https://api.typesafe.ai/v1/systemone, and the body below. Column references like /Company Name must match your table and stay inside quotes.

6

Pick the fields to return

answers.icp_fit.score, answers.icp_fit.confidence, answers.is_decision_maker.noul and answers.is_competitor.noul.

7

Rate-limit it

About 10 requests per 1,000 ms leaves headroom under TypeSafe’s 1,200 a minute. Add an Only run if condition so rows without a description are skipped.

8

Run one row, then the table

Check the first result before spending on the rest.

9

Clean up the score

A five-level score comes back as 0 to 4 and can land between levels, so add a formula column that adds 1 and rounds it for a clean 1 to 5. Flag rows with confidence under 0.5 for a quick human look.

10

Push to HubSpot

Use Clay’s HubSpot actions: Look up object by email first, then Create object for new contacts and Update object for existing ones. Create the custom properties icp_score and icp_confidence in HubSpot first so they show up in the mapping.

The request body

Replace the ideal-customer sentence with who you sell to and the problem you solve.

02-lead-fit.json31 lines
{
  "model": "jev-1.13.0",
  "state": {
    "company": "/Company Name",
    "industry": "/Industry",
    "company_description": "/Company Description",
    "job_title": "/Title",
    "headline": "/LinkedIn Headline"
  },
  "questions": {
    "icp_fit": {
      "type": "score",
      "instructions": "How well do `company`, `company_description` and `job_title` match our ideal customer: REPLACE with one sentence describing who you sell to and the problem you solve.",
      "criteria": [
        "Clearly outside our market: consumer business, student, job seeker, or a company that sells the same service we do",
        "Wrong industry, with no sign they sell to other businesses",
        "Sells to businesses, but the person is junior or not involved in sales or operations",
        "Right kind of business and a manager-level buyer, but no visible sign of the problem we solve",
        "Right kind of business, a founder, owner or head of sales or operations, and visible signs they have the problem we solve"
      ]
    },
    "is_decision_maker": {
      "type": "noul",
      "instructions": "`job_title` is a founder, owner, C-level, VP or head of sales, growth or operations."
    },
    "is_competitor": {
      "type": "noul",
      "instructions": "`company_description` describes a business that sells the same service we do."
    }
  }
}

Two upgrades. Filter headcount, revenue and funding in a Clay formula before or after the Jev call, never inside the question. And once it works, split the one big fit score into small questions (industry fit, buyer role, visible need) and combine them with weights in a formula. TypeSafe calls this composite scoring, and it is more reliable than one giant question.

Check this in your account: run a row whose description contains quote marks. If Clay returns a body parse error, clean the column with a formula first.

What it costs: about 600 input tokens a lead. 100,000 leads come to around 60 million tokens, about $2.52 in Jev. Clay’s credits for the HTTP column are billed separately.

Use 3 of 7 · Customer

Sort 10,000 replies for 34 cents

Every reply labelled in about a third of a second. By the guide’s own label set, “Free Thursday at 2?” would be book_meeting rather than interested.

Every reply, email or support ticket gets a label in about a third of a second: interested, not now, unsubscribe, billing and so on. The hot ones ping your team in Slack, and nobody reads 300 out-of-office replies to find the two people who said yes.

You need Zapier, the TypeSafe Jev app, whichever inbox you use (Gmail, Crisp or Instantly), and Slack.

The Zap

1

Trigger

Gmail New Email (or New Labeled Email), Crisp Message Incoming, or Instantly New Event with the event type Reply Received. Instantly also has native webhooks on its Hyper Growth plan, but the Zapier trigger is the simpler route.

2

TypeSafe Jev › Ask Questions

Content: the email or message body from the trigger, up to about 24,000 words. Pick one, name: reply_type, question: “What kind of reply is this?”, options: interested, book_meeting, referral, not_now, objection, not_interested, unsubscribe, unclear. Yes-or-no questions: “The reply asks a question that needs an answer from us.” and “The reply mentions a competitor or current provider.” Minimum confidence: 0.6 to start.

3

Filter

Filter by Zapier, or Paths, on the Jev answer. Let interested and book_meeting through.

4

Notify

Slack › Send Channel Message with the sender, the label and a link to the thread. Optional extras: Gmail Add Label to Email, or Crisp Add Segments to Conversation.

The labels, with the descriptions that make Jev accurate

For support tickets, swap these for your queues (billing, technical, account access, feature request, cancellation, other) and describe each in a sentence.

LabelMeans
interestedWants to learn more, asks a buying question, or says yes to a call
book_meetingProposes or accepts a specific time, or asks for a calendar link
referralNot the right person but names or forwards to someone else
not_nowInterested in principle but says later, next quarter, or after an event
objectionPushes back on price, fit, or current vendor but stays in the conversation
not_interestedClear no, with no request to stop emailing
unsubscribeAsks to stop emailing, remove them, or threatens to report spam
unclearNone of the above fits, or the reply is too short to tell

The same question as raw JSON

For the API, Clay, ManyChat or your own code.

03-reply-labels.json22 lines
{
  "model": "jev-1.13.0",
  "state": { "reply": "PASTE THE REPLY TEXT" },
  "questions": {
    "reply_type": {
      "type": "choice",
      "instructions": "What kind of reply is `reply`?",
      "criteria": {
        "interested": "Wants to learn more, asks a buying question, or says yes to a call",
        "book_meeting": "Proposes or accepts a specific time, or asks for a calendar link",
        "referral": "Not the right person but names or forwards to someone else",
        "not_now": "Interested in principle but says later, next quarter, or after an event",
        "objection": "Pushes back on price, fit, or current vendor but stays in the conversation",
        "not_interested": "Clear no, with no request to stop emailing",
        "unsubscribe": "Asks to stop emailing, remove them, or threatens to report spam",
        "unclear": "None of the above fits, or the reply is too short to tell"
      }
    },
    "needs_answer": { "type": "noul", "instructions": "`reply` asks a question that needs an answer from us." },
    "mentions_competitor": { "type": "noul", "instructions": "`reply` mentions a competitor or current provider." }
  }
}

Check this in your account: the Zapier app is only days old, so the first time you set it up, check how the options field wants its list, commas or one per line. Two behaviours worth knowing. Answers below your minimum confidence come back as “unsure”, and the only choice for those is Stop the Zap, which ends the run quietly like a Filter. And a pick-one question answers “none” when no option fits. The action also takes a scale question, with 2 to 10 levels listed lowest first.

What it costs: about 800 input tokens an email with the question. 10,000 replies come to around 8 million tokens, about 34 cents. Zapier tasks are billed separately.

Use 4 of 7 · Ops

Stop your AI agent before it does something stupid

Every action gets a yes or no before it runs, and the risky ones are blocked.

Before your agent runs a tool, Jev asks one question: how likely is this call to be risky or not properly authorised? Safe calls run and risky ones are stopped. Reading a calendar goes through; emailing 4,212 contacts doesn’t.

If your agent runs on LangChain

install.sh2 lines
pip install "langchain-typesafe[experimental]" langchain-openai
export TYPESAFE_API_KEY=...

The guard

From the starter kit. Only the tools you list get checked, so list the dangerous ones.

04-agent-guard.py35 lines
# Use case 4: stop risky tool calls before they run.
# pip install "langchain-typesafe[experimental]" langchain-openai
# export TYPESAFE_API_KEY=...   (and your model provider key)
# Source: https://docs.langchain.com/oss/python/integrations/providers/typesafe
#
# AutoModeMiddleware asks Jev how likely each listed tool call is to be risky or not properly authorised.
# Risky calls return an error message to the agent instead of running. It blocks; it does not ask you for approval.
# It is experimental. Keep hard permissions in code as well: text the agent reads can still sway the answer.

from langchain.agents import create_agent
from langchain.tools import tool
from langchain_typesafe.experimental.middleware import AutoModeMiddleware


@tool
def send_email_to_list(list_name: str, subject: str) -> str:
    """Send an email to every contact on a list."""
    return f"Sent '{subject}' to {list_name}."


@tool
def read_calendar(day: str) -> str:
    """Read the calendar for a day."""
    return f"Calendar for {day}: 2 meetings."


agent = create_agent(
    "openai:gpt-6-astra",
    tools=[send_email_to_list, read_calendar],
    # Only the tools you list here get checked. Put your dangerous ones in.
    middleware=[AutoModeMiddleware(tools=[send_email_to_list])],
)

result = agent.invoke({"messages": [{"role": "user", "content": "Email the whole newsletter list about the price rise."}]})
print(result["messages"][-1].content)

A blocked call returns an error message to the agent instead of running. The middleware blocks rather than asking for approval, so add LangChain’s human-in-the-loop middleware, which needs a checkpointer, if you want a person to sign off. The guard itself is experimental and fixed in two ways. It blocks any call it rates 0.5 or more likely to be risky, and you can’t change that threshold. And if the check itself fails, the tool doesn’t run. It sends the tool call and up to 30 recent messages to TypeSafe, so keep secrets out of tool arguments.

If you build on Vercel’s eve framework, which is in beta, its auto() approval helper asks Jev to sort each tool call into “clear” or “caution”. Clear runs, and caution pauses for a person. It sees only the tool’s name and arguments, not the conversation.

Any other stack can use the pattern from OpenRouter’s cookbook. Run the plain checks in code first, such as a refund amount against a balance, then ask Jev three yes-or-no questions about the call: did the customer ask for this, is it the right record, does the policy cover it. Approve when all three are at 0.9 or above, block when any is at 0.1 or below, and send everything in between to a person.

!
Read this before you ship it

TypeSafe’s own docs say text written to steer the model can move the answer, and that includes instructions hidden in an email or a web page your agent reads. So treat Jev as a second check, not the lock. Keep real permissions in code: an agent with no access to the delete button can’t press it.

What it costs: a few hundred tokens a check. OpenRouter’s cookbook measured each request, with all three questions in it, at under $0.0001 and under 600 ms.

Use 5 of 7 · Marketing

Check everything your AI writes

Before a draft goes out, Jev asks the questions you would. The slide’s “sounds like our brand?” isn’t one of the kit’s six checks; the nearest is “sounds like marketing copy”.

Before a draft goes out, Jev asks the questions you would. Does it make a claim that isn’t in your notes? Does it sound like an ad? Does it name a competitor? Anything flagged gets sent back or pinged to you.

No-code version, in Zapier

1

Trigger

Wherever your drafts live: a Google Doc, an Airtable record, a form you paste into.

2

TypeSafe Jev › Ask Questions

For Content, send the facts and the draft together, like “SOURCE NOTES: … DRAFT: …”. The fact checks need something to compare against.

3

Add the checks

The yes-or-no checks below, with a minimum confidence around 0.7.

4

Filter

Filter by Zapier: carry on only if any check came back yes.

5

Notify

Slack › Send Channel Message to yourself, with the check that fired.

If your team already logs AI calls in Langfuse, it has a no-code “Jev as a judge” evaluator. Add an LLM connection with the typesafe adapter, create a New decision model evaluator, write the questions, map the input and output fields, test it and attach it to a rule. Every question writes a score you can chart.

Checks worth running

✓The draft states a number, result, client name or promise that isn’t in the source notes✓The draft sounds like marketing copy rather than a person writing to one person✓The draft opens with a pleasantry or filler before making its point✓The draft makes a guarantee or claims certainty about results✓The draft names a competitor or another company’s product✓The draft asks for a meeting more than once

The six checks as one request

From the starter kit. Replace the two REPLACE lines, or map them to your fields.

05-draft-checks.json15 lines
{
  "model": "jev-1.13.0",
  "state": {
    "source_notes": "REPLACE: the facts the draft is allowed to use (results, names, numbers, offer details)",
    "draft": "REPLACE: the draft your AI wrote"
  },
  "questions": {
    "unsupported_claim": { "type": "noul", "instructions": "`draft` states a number, result, client name or promise that does not appear in `source_notes`." },
    "sounds_like_marketing": { "type": "noul", "instructions": "`draft` sounds like marketing copy rather than a person writing to one person." },
    "filler_opening": { "type": "noul", "instructions": "`draft` opens with a pleasantry or filler before making its point." },
    "guarantees_results": { "type": "noul", "instructions": "`draft` makes a guarantee or claims certainty about results." },
    "names_competitor": { "type": "noul", "instructions": "`draft` names a competitor or another company's product." },
    "pushes_meeting_twice": { "type": "noul", "instructions": "`draft` asks for a meeting or call more than once." }
  }
}

Keep exact checks out of Jev: word counts, em dashes, the number of links. A formatter step or two lines of code does those perfectly.

How good is it? In Every’s test it caught 6 of 7 planted mistakes, and Fable 5.1 caught all 7. Langfuse cites a larger study in which Jev agreed with Fable 5.1’s verdict 91.5% of the time across 6,003 checks, at about $160 per million graded answers against $33,000 for Fable. That is agreement rather than accuracy, and a cheap chat model, DeepSeek V4.1 Flash, agreed slightly more often. Use Jev as the first pass that catches most problems cheaply; anything that really matters still gets a human read.

What it costs: about 1,100 input tokens a draft with six checks. 1,000 drafts cost about 5 cents.

Use 6 of 7 · Intel

Find out why your best posts pop

Example data, as the slide says. It mentions hook, topic and format; the kit’s script tags hook, format and call to action.

Jev tags every post you’ve made by hook, format and call to action. Put the tags next to your views, sort, and you’ll see which hooks and formats pull the views.

Export your posts

1

Instagram

Meta Business Suite › Insights › Content › Export. Menu names shift between accounts, so look for the export button on the content view.

2

YouTube

YouTube Studio › Analytics › Advanced mode › Export current view. Downloaded reports stop at 500 rows.

3

Add transcripts if you have them

Jev reads text only, so a transcript column helps a lot.

4

Tidy the file

Four columns: id, caption, transcript, views.

Run the script

It writes tagged.csv and prints the median views for every hook type, format and call to action. No installs needed: it uses Python’s standard library only.

terminal3 lines
export OPENROUTER_API_KEY=...        # or TYPESAFE_API_KEY
python3 06-tag_posts.py posts.csv --dry-run   # shows the first request, sends nothing
python3 06-tag_posts.py posts.csv            # tags every post

The tags it uses

Edit 06-post-tags.json to match your content, and keep an “other” option in every pick-one question.

TagTypeOptions
hook_typeTypePick oneOptionsnumber claim, tool reveal, problem callout, contrarian, question, story, other
formatTypePick oneOptionstutorial, screen demo, build reveal, list of tools, case study, opinion, other
cta_typeTypePick oneOptionscomment keyword, link in bio, follow, DM, none
names_a_toolTypeYes or noOptionsThe caption or transcript names a specific AI tool
shows_resultTypeYes or noOptionsThe post shows a finished result

The tag questions

06-post-tags.json41 lines
{
  "hook_type": {
    "type": "choice",
    "instructions": "What kind of opening does the post use? Judge the first line of `caption` and the first sentence of `transcript`.",
    "criteria": {
      "number_claim": "Opens with a specific result, number, or dollar figure",
      "tool_reveal": "Opens by naming a tool or showing something built",
      "problem_callout": "Opens by naming a pain the viewer has",
      "contrarian": "Opens by disagreeing with common advice",
      "question": "Opens with a question to the viewer",
      "story": "Opens with a personal story or moment",
      "other": "None of the above"
    }
  },
  "format": {
    "type": "choice",
    "instructions": "What format is the post, judging by `caption` and `transcript`?",
    "criteria": {
      "tutorial": "Talks the viewer through how to do something",
      "screen_demo": "Shows a screen recording of a tool or build",
      "build_reveal": "Shows a finished build and what it does",
      "list_of_tools": "Runs through a list of tools, tips or resources",
      "case_study": "Tells what happened for a client or project",
      "opinion": "Gives a take or prediction",
      "other": "None of the above"
    }
  },
  "cta_type": {
    "type": "choice",
    "instructions": "What does `caption` ask the viewer to do?",
    "criteria": {
      "comment_keyword": "Comment a word to get something",
      "link_in_bio": "Go to the link in bio",
      "follow": "Follow the account",
      "dm": "Send a direct message",
      "none": "No ask"
    }
  },
  "names_a_tool": { "type": "noul", "instructions": "`caption` or `transcript` names a specific AI tool or product." },
  "shows_result": { "type": "noul", "instructions": "The post shows a finished result or output, going beyond an explanation." }
}

The tagging script

Checked here: with --dry-run it prints the first request and sends nothing, and fed a stubbed reply it writes tagged.csv and the median-views summary. The 0.06-second pause keeps it under 1,200 requests a minute.

06-tag_posts.py122 lines
#!/usr/bin/env python3
"""Use case 6: tag every post with Jev, then see which kinds of post get the views.

Input:  a CSV with columns id, caption, transcript (optional), views
        (export from Meta Business Suite or YouTube Studio, then tidy the column names)
Output: tagged.csv (your columns plus one column per tag) and a summary of median views per tag

Set ONE key first:
  export TYPESAFE_API_KEY=...     calls https://api.typesafe.ai/v1/systemone
  export OPENROUTER_API_KEY=...   calls https://openrouter.ai/api/alpha/decisions

Run:
  python3 06-tag_posts.py sample-posts.csv --dry-run   print the first request, send nothing
  python3 06-tag_posts.py posts.csv                    tag everything

No installs needed (standard library only). The questions live in 06-post-tags.json: edit the
options to match your content, and keep an "other" option in every pick-one question.
"""
import csv
import json
import os
import statistics
import sys
import time
import urllib.error
import urllib.request
from collections import defaultdict
from pathlib import Path

HERE = Path(__file__).parent
QUESTIONS = json.loads((HERE / "06-post-tags.json").read_text())
MAX_TRANSCRIPT_CHARS = 6000  # the opening matters most; this also keeps each call cheap


def endpoint():
    if os.environ.get("TYPESAFE_API_KEY"):
        return "https://api.typesafe.ai/v1/systemone", os.environ["TYPESAFE_API_KEY"], "jev-1.13.0"
    if os.environ.get("OPENROUTER_API_KEY"):
        return "https://openrouter.ai/api/alpha/decisions", os.environ["OPENROUTER_API_KEY"], "typesafe/jev-1.13"
    return None, None, "jev-1.13.0"


def build_body(model, post):
    state = {"caption": (post.get("caption") or "").strip(),
             "transcript": (post.get("transcript") or "").strip()[:MAX_TRANSCRIPT_CHARS]}
    return {"model": model, "state": state, "questions": QUESTIONS}


def ask(url, key, body):
    data = json.dumps(body).encode()
    req = urllib.request.Request(url, data=data, headers={"Authorization": f"Bearer {key}", "Content-Type": "application/json"})
    for attempt in range(5):
        try:
            with urllib.request.urlopen(req, timeout=30) as r:
                return json.loads(r.read())
        except urllib.error.HTTPError as e:
            if e.code in (429, 500, 502, 503, 529) and attempt < 4:
                time.sleep(2 ** attempt)  # rate limited or busy: back off and retry
                continue
            raise


def flatten(answers):
    row = {}
    for name, a in answers.items():
        kind = a.get("type")
        if kind == "choice":
            row[name] = a.get("choice")
            row[name + "_confidence"] = round(a.get("confidence", 0), 3)
        elif kind == "score":
            row[name] = round(a.get("score", 0), 2)
            row[name + "_confidence"] = round(a.get("confidence", 0), 3)
        elif kind == "noul":
            row[name] = round(a.get("noul", 0), 3)  # probability of yes
    return row


def to_number(v):
    try:
        return float(str(v).replace(",", "").strip())
    except ValueError:
        return None


def summary(rows, tag):
    groups = defaultdict(list)
    for r in rows:
        views = to_number(r.get("views"))
        if views is not None and r.get(tag):
            groups[r[tag]].append(views)
    print(f"\nMedian views by {tag}:")
    for value, views in sorted(groups.items(), key=lambda kv: -statistics.median(kv[1])):
        print(f"  {value:<20} {statistics.median(views):>12,.0f}   ({len(views)} posts)")


def main():
    if len(sys.argv) < 2:
        sys.exit(__doc__)
    posts = list(csv.DictReader(open(sys.argv[1], newline="", encoding="utf-8-sig")))
    url, key, model = endpoint()
    if "--dry-run" in sys.argv:
        print(json.dumps(build_body(model, posts[0]), indent=2))
        return
    if not key:
        sys.exit("Set TYPESAFE_API_KEY or OPENROUTER_API_KEY first.")
    out = []
    for i, post in enumerate(posts, 1):
        res = ask(url, key, build_body(model, post))
        out.append({**post, **flatten(res.get("answers", {}))})
        print(f"{i}/{len(posts)} {post.get('id', '')}: {out[-1].get('hook_type')}")
        time.sleep(0.06)  # stays under 1,200 requests a minute
    with open("tagged.csv", "w", newline="", encoding="utf-8") as f:
        w = csv.DictWriter(f, fieldnames=list(out[0].keys()))
        w.writeheader()
        w.writerows(out)
    print("\nWrote tagged.csv")
    for tag in ("hook_type", "format", "cta_type"):
        summary(out, tag)


if __name__ == "__main__":
    main()

Three rows to try it on

sample-posts.csv4 lines
id,caption,transcript,views
p1,"This AI model is 200x faster and 400x cheaper. Here are the 7 best things to use it for.","",412000
p2,"Stop doing your follow-ups by hand. Comment FLOW and I'll send you the setup.","If you're still following up by hand, you're leaving deals on the table. Here's the system I use.",188000
p3,"What would you build with a second brain?","So I asked my AI a question.",64000

A tip from OpenRouter’s own tests: if a tag keeps coming back wrong, the fix is almost always the question. A tag whose description overlaps another tag never gets accurate, whatever threshold you set. Rewrite the description and run it again; it costs pennies.

What it costs: about 1,900 input tokens a post with a transcript. 1,000 posts cost about 8 cents.

Use 7 of 7 · Back office

Stop paying top-model prices for easy jobs

Most of what an AI app does is simple. Jev spots the easy jobs and sends them to a cheaper model.

Most of what your AI app handles is simple: order lookups, summaries, translations. Jev reads each request and sends the easy ones to a cheap model and the hard ones to the expensive one.

If you run a LiteLLM proxy, its Auto Router can use Jev as the classifier. Point your app at jev-router and set TYPESAFE_API_KEY on the proxy. The tier names must match models defined in the same config.

The LiteLLM config

The full file from the starter kit, tiers included. Swap in the models you actually use.

07-litellm-router.yaml35 lines
# LiteLLM proxy config: Jev decides which tier each request needs.
# Needs TYPESAFE_API_KEY in the proxy environment. Tier values must be model_name entries defined elsewhere in this file.
# Source: https://docs.litellm.ai/docs/auto_router/setup
model_list:
  # The four tiers. Swap these for the models you actually use; keep the model_name values matching the tiers below.
  - model_name: gpt-5.6-luna
    litellm_params: { model: openai/gpt-5.6-luna, api_key: os.environ/OPENAI_API_KEY }
  - model_name: gpt-5.6-terra
    litellm_params: { model: openai/gpt-5.6-terra, api_key: os.environ/OPENAI_API_KEY }
  - model_name: claude-sonnet-5
    litellm_params: { model: anthropic/claude-sonnet-5, api_key: os.environ/ANTHROPIC_API_KEY }
  - model_name: claude-opus-5
    litellm_params: { model: anthropic/claude-opus-5-5, api_key: os.environ/ANTHROPIC_API_KEY }

  # The router your app calls. Point your app at model "jev-router".
  - model_name: jev-router
    litellm_params:
      model: auto_router/complexity_router
      complexity_router_default_model: gpt-5.6-terra
      complexity_router_config:
        tiers:
          SIMPLE: gpt-5.6-luna
          MEDIUM: gpt-5.6-terra
          COMPLEX: claude-sonnet-5
          REASONING: claude-opus-5
        classifier_type: jev
        jev_classifier_config:
          model: jev-latest
          timeout_ms: 3000
          circuit_breaker_enabled: true
          circuit_breaker_cooldown_seconds: 30
        classifier_fallback: default_model
        classifier_context_window_size: 3
        classifier_context_budget_chars: 8000
        classifier_context_include_assistant_turns: false

In LiteLLM’s own benchmark, Jev picked a tier in about 127 ms against 688 ms for Claude Haiku 4.5 doing the same job, at 96% lower classifier cost, and matched LiteLLM’s intended tiers 95% of the time against 74% for Haiku. LiteLLM wrote the 80 test prompts and their expected tiers itself, and it measured the routing, not the answers. Two defaults to know. A failed classification falls back to LiteLLM’s own heuristic unless you set classifier_fallback: default_model, as this file does. And earlier user turns are sent to TypeSafe for context.

If you build with LangChain, ModelRouterMiddleware does the same inside an agent: you describe each model’s job in a sentence, and Jev picks one per run.

The LangChain version

07-model-router.py24 lines
# Use case 7 (LangChain version): Jev reads the request and picks the cheapest model that can do it.
# pip install "langchain-typesafe[experimental]" langchain-openai
# Source: https://docs.langchain.com/oss/python/integrations/providers/typesafe

from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import ModelChoice, ModelRouterMiddleware

router = ModelRouterMiddleware(
    choices={
        "fast": ModelChoice(
            model="openai:gpt-5.6-terra",
            criteria="Direct lookups, extraction, and localized changes with explicit targets.",
        ),
        "powerful": ModelChoice(
            model="openai:gpt-6-astra",
            criteria="Architecture, novel root-cause reasoning, and high-stakes decisions.",
        ),
    },
    instructions="Choose the least costly model that can complete the task safely.",
)

agent = create_agent("openai:gpt-5.6-terra", middleware=[router])
result = agent.invoke({"messages": [{"role": "user", "content": "What's the status of order 4471?"}]})
print(result["model_route"])  # which model Jev picked, with probabilities and confidence

The cascade pattern, from OpenRouter’s cookbook, goes one step further. A cheap model drafts the answer, Jev checks whether the draft is supported, and only the failures go to the big model. The guide reports that on OpenRouter’s 50-question test the cascade got every answer right for $0.012, while the big model alone got 2 wrong for $0.175. The cookbook itself is less flattering. The cheap model on its own also made no mistakes, for $0.004. The cascade passed 2 answerable questions to a person. And OpenRouter calls a two-answer gap between runs noise. Try it on your own traffic before you count on it.

What it costs: even with a few turns of conversation included, a routing check is a couple of thousand tokens at most, so under 10 cents per 1,000 requests. What you save depends on how much of your traffic is easy, and a day of your logs will tell you.

The cost math

Every number above, from one formula

Input tokens per item × items ÷ 1,000,000 × $0.042. Output is free. The question counts too, because it is sent with every call. Other tools bill on their own: Clay credits, Zapier tasks, your ManyChat plan.

UseTokens per item, with the questionVolumeJev cost
Comments and DMsTokens per item, with the questionabout 400Volume10,000Jev costabout $0.17
Lead scoringTokens per item, with the questionabout 600Volume100,000Jev costabout $2.52
Replies and ticketsTokens per item, with the questionabout 800Volume10,000Jev costabout $0.34
Agent checksTokens per item, with the questiona few hundredVolume10,000Jev costunder $1
Draft checksTokens per item, with the questionabout 1,100Volume1,000Jev costabout $0.05
Post taggingTokens per item, with the questionabout 1,900Volume1,000Jev costabout $0.08
Model routingTokens per item, with the questionup to 2,000Volume1,000Jev costunder $0.10
Limits

Where Jev is the wrong tool

✗Anything that needs writing. It never writes a sentence, so pair it with a writing model✗Counting and maths. It can’t count words or compare numbers reliably; do that in code✗Images, audio and video. It reads text only, so use captions and transcripts✗Final calls on things that matter. It catches most problems, not all, so keep a human on the ones that cost real money✗Hostile text. Text written to steer it can move the answer, so keep real permissions in code✗Perfectly repeatable scores. OpenRouter measured confidence moving by up to 0.13 between identical runs, so leave a margin around your thresholds✗Very long documents. TypeSafe allows 64,000 tokens a request but only 32,000 for the text plus the longest question, and some providers list 32K in total. Split long documents first

TypeSafe’s published evals score Jev against the averaged answers of GPT-6 Astra and Claude Fable 5.1, both at high thinking. Its lowest scores are security incidents at 61.7% and invoices at 61.8%, against 79.1% for OpenAI’s Sol on invoices, and its mean across all of them is 67.8%. TypeSafe’s own list of weak spots adds literal readings of questions, date and time comparisons, questions that need several steps of reasoning, and long runs of irrelevant text. Related questions don’t have to add up, either: two opposite yes-or-no questions once summed to 1.19.

Test it on 50 of your own examples before you trust it with a whole pipeline. In TypeSafe’s words, “Jev guarantees the shape of its answers, not that every decision is correct.”

The starter kit

Every file, and where it goes

All of them are on this page, in the sections above.

FileUseWhere it goes
first-call.shUseTest your keyWhere it goesTerminal
01-buyer-intent-manychat-body.txtUse1. Buyers in comments and DMsWhere it goesManyChat, Make External Request, Body. Replace the offer line
02-lead-fit.jsonUse2. Lead scoringWhere it goesClay, HTTP API column, Body. Column names must match your table
03-reply-labels.jsonUse3. Sorting replies and ticketsWhere it goesThe Zapier fields are in the steps; this is the raw JSON version
04-agent-guard.pyUse4. Blocking risky agent actionsWhere it goesLangChain agent code
05-draft-checks.jsonUse5. Checking AI draftsWhere it goesZapier, Langfuse or your own code
06-post-tags.json, 06-tag_posts.py, sample-posts.csvUse6. What makes posts workWhere it goesTerminal: python3 06-tag_posts.py posts.csv
07-litellm-router.yaml, 07-model-router.pyUse7. Easy jobs to cheap modelsWhere it goesLiteLLM proxy config, or LangChain
Before you switch it on

Checklist

✓A key from TypeSafe or OpenRouter, or the Zapier app connected✓first-call.sh returns an answer✓One use case picked to start with; comments and DMs is the fastest win✓Every option described in a sentence, with an “other” or “unclear” option✓Numbers and counts handled in a formula or code, outside Jev✓Thresholds set: act above 0.8, and a person looks at the middle✓Tested on 20 to 50 real examples before switching it on✓Model pinned to jev-1.13.0 once it works✓Real permissions kept in code for anything an agent can do