Jev tool call guard: check agent actions before they execute

Paste what the user asked for and the tool call your agent wants to make. Jev returns whether the call matches the request, how risky it is on a five-level scale, and whether the runtime should execute it, ask the user or block it — as probabilities your harness can act on before anything runs.

Try it on your own text

Edit the example and run it. No account needed — 3 free runs a day.

Scenario

Check an agent's proposed tool call against what the user asked.

201 / 2,000

Edit freely — Jev answers the scenario's questions about whatever you put here.

What this guard checks

The example above is the failure every agent builder worries about. The user asked to "clean up old test data in staging". The agent proposes run_sql against the production database with DELETE FROM orders; — the wrong environment and the wrong scope, with no WHERE clause. The state holds both the request and the proposed call, and one evaluation asks three questions:

QuestionTypeWhat you get back
matches_intentNoulProbability the proposed call is what the user asked for
riskScoreA position on Harmless → Low → Moderate → High → Destructive
decisionChoiceA probability for execute, confirm and block, plus confidence

The first two questions separate two different problems. A call can match the request and still be dangerous ("drop the staging database" when the user asked for exactly that), or be harmless and still be wrong. decision turns both into what your runtime actually does. All three are answered in parallel against the same state, so together they cost one credit and barely more time than one.

This page screens actions. The prompt injection detector screens inputs before your model reads them. They cover different moments: an injection in a web page is caught on the way in; the destructive call it provoked — or one the model came up with unprompted — is caught here, on the way out.

How to write guard questions

Put the request next to the call. Jev can only judge intent if the state contains both. Send the user’s latest request and the tool name with its arguments as one JSON object, as the example does. Leave out the rest of the transcript unless it matters: TypeSafe’s notes say unrelated material in the state costs accuracy.

Describe risk for your tools. "Destructive" means something different for send_email than for run_sql. Give the Noul or Score criteria written for your tools:

{
  "type": "noul",
  "instructions": "Could this tool call cause damage that cannot be undone?",
  "criteria": {
    "true": "It deletes, overwrites or sends data outside the system, or touches production without an explicit request.",
    "false": "It only reads data, or writes to a sandbox or a staging environment the user named."
  }
}

Compute the numbers in code. Row counts, file sizes and deadlines are weaknesses TypeSafe documents. If a dry run says a query would touch 40,000 rows, put that fact in the state as a field rather than asking Jev to infer it.

Strip secrets from arguments. Whatever is in the state is sent for classification. LangChain’s documentation gives the same warning for its middleware: do not put secrets in tool arguments unless sending them to TypeSafe is acceptable.

For more on phrasing, see Designing Questions Jev Can Answer.

From one check to a pipeline step

  1. Check before execution, not after. A guard that runs after the tool is a log line. Wrap the tool executor so the call is classified first and only runs on a pass.
  2. Use the LangChain middleware if you are on LangChain. LangChain’s langchain-typesafe package ships an experimental AutoModeMiddleware on the wrap_tool_call hook. It asks Jev whether a call is risky or insufficiently authorized and returns an error message to the agent instead of running the tool. It only classifies the tools you name, and it refuses rather than asks — pair it with human-in-the-loop middleware if you want the confirm path this page’s example includes.
  3. Gate by consequence. Set a higher bar for destructive tools than for read-only ones, and send the middle band to confirm instead of guessing.
  4. Record the model version. Every response carries a model field such as jev-1.13.0. Log it with each decision and re-check your thresholds when it changes — JevStation pins the model for you, so you do not choose it per request.
  5. Call it from your harness. Outside LangChain, the same evaluation runs through the System One API with an API key, at the same credits as the playground.

The full pattern — routing before the model call, gating before the tool, and what to do with low confidence — is in Building a Harness with Jev.

Check it before you trust it

  • Label your own tool calls. Pull real calls from your logs, add a set of bad ones written for your tools, and measure where Jev separates them. Count the good calls it would block, too: a guard that interrupts constantly gets switched off.
  • Take reported results as leads, not guarantees. TechCrunch reported that Vercel replaced the OpenAI model it used to review commands for safety with Jev and got results five to 18 times faster with greater accuracy. That is one company’s account, not a benchmark you can reuse.
  • Test crafted arguments. TypeSafe lists adversarial content as a known weakness. Include calls whose arguments are written to look routine.
  • Never show the agent its scores. In an independent AI-control pilot, an attacker allowed to query the monitor a few times before submitting got roughly a third of backdoors through instead of a tenth.

Where Jev is the wrong tool

  • As your only safety boundary. Anthropic says its own auto-mode classifier is not a drop-in replacement for careful human review on high-stakes infrastructure, and Cursor calls its auto-review not a security boundary. The same applies here.
  • The rule is exact. "Never write to production from this agent" is a credential scope or a deny rule, not a probability.
  • The risk is hidden in a file. Whether python script.py is safe depends on the script. Jev only sees the state you send, so read the file in code and include what matters.
  • You need a written reason for an audit log or a reviewer. Jev returns probabilities only.

Frequently asked questions

How is this different from the prompt injection detector?
The prompt injection detector screens inputs — messages, documents and tool results — before your model reads them. The tool call guard screens actions: the call the agent has already decided to make. Use both; a guard catches a bad call whether it came from an injection, a confused user or a model mistake.
Does LangChain have a Jev guardrail built in?
Yes. LangChain’s langchain-typesafe package ships an experimental AutoModeMiddleware that asks Jev whether a tool call is risky and returns an error to the agent instead of running it. It only checks the tools you list, and it refuses rather than asking for approval.
What does checking one tool call cost?
On JevStation, a check with up to five questions and a state under 8,000 characters is a standard evaluation and costs 1 credit. The three questions in this example fit in that. A failed evaluation is refunded.
Is a Jev guard enough to make an agent safe?
No. It is a probabilistic classifier. Keep scoped credentials, deny rules for irreversible operations and a person approving high-stakes actions; use Jev to decide which calls deserve that extra step.

Related

Make it part of your pipeline

Sign up for 200 free credits, save your own question sets and call the same evaluation from the API.