Writing, 30 Sept 2026, 3 min read

jevbrief: give an AI model less noise and get better decisions

jevbrief trims logs, tools and web pages down to what matters before an AI model decides, and records why every item was dropped. With benchmarks.

AI models that make decisions do their best work on small, relevant input. Real sources are the opposite: a failed build prints thousands of log lines, an agent can call a hundred tools, a web page has hundreds of elements. I built jevbrief, an open-source Python library, to sit between the source and the model: it keeps what matters, drops the rest with a reason for every drop, and records the whole decision so you can replay it.

The problem

jevbrief is built for Jev, TypeSafe’s model that turns state into typed decisions, like “which error broke this build” or “which tool should the agent call next”. Jev is strong at judgment and weak at raw numbers, dates and long lists of irrelevant detail.

So the naive integration fails in a predictable way. You pipe 1,395 lines of CI logs into the model, about 29,700 tokens, and ask what broke. It is slow, expensive, and the real error is buried in noise. Worse, when the answer is wrong, you can’t tell what the model was actually shown.

Other tools show what the model decided. jevbrief shows what it was told, what it wasn’t told, and why.

What it does

  • Cuts input down to what matters. In one example, 1,395 log lines at 29,678 tokens become 5 log groups at 1,070 tokens.
  • Explains every drop. Each removed item gets a reason code, so nothing disappears silently.
  • Asks one clear question. The model gets a focused goal, not a wall of text.
  • Records the decision. Every call is traced, and --view replays it.
  • Works on many sources. Adapters cover CI logs, pull requests, OpenTelemetry signals, agent tools, agent step histories, JSON, web pages, Android screens, security alerts and even NES games.

How it works

Each adapter reads a source and turns it into a small set of facts. Rules filter those facts, dropping noise with a reason code. What survives goes to Jev with one question, and Jev picks an answer with a confidence score.

For agent builders, it also works as a library:

from jevbrief import select_tools, pick_tool, check_progress

llm.bind_tools(select_tools(tools, goal))   # narrow 100 tools to the relevant few, locally, free
pick = pick_tool(tools, goal)               # Jev picks one: pick.tool, pick.confidence
check_progress(history, goal).stuck         # is the agent going in circles?

The tools can be your own functions, MCP servers, or LangChain, CrewAI, OpenAI or Anthropic tools, and you get your own objects back.

Try it

pip install jevbrief
gh run view <run id> --log-failed > run.log
jevbrief inspect run.log --adapter ci --goal "CI is red on main"

inspect needs no API key and costs nothing: it shows what would be sent. ask calls Jev with a TYPESAFE_API_KEY and returns an answer like:

27 log lines -> 7 failures -> 3 sent to Jev
likely cause: Run pytest -q: E KeyError: 'currency_code'   (confidence 0.84)

What the benchmarks showed, including where it lost

I compared jevbrief against a naive integration that sends the raw source, with the same model, the same question, and three runs per task.

Source Accuracy, raw to jevbrief Input tokens, raw to jevbrief
CI logs, 16 real failed runs 88% to 100% 30,502 to 988
Agent tools, 105 real MCP tools 82% to 88% 13,450 to 3,685
Android taps, 105 real taps 61% to 84% 18,753 to 2,049

Not every result went my way, and I think that part is worth publishing too:

  • Flaky tests: on whether a CI failure looks flaky, jevbrief scored 67% against 93% for the raw logs. Trimming removed exactly the context that question needs.
  • Pull requests: jevbrief tied the full diff on accuracy. It only saved tokens.
  • Security alerts: jevbrief scored below the raw alerts, and below simply answering “not exploitable” every time. It found far more of the truly exploitable cases, 76% against 32%, but with twice the false alarms.
  • Synthetic data: some test sets were built alongside the adapters, which flatters them.

Two bugs also taught me something. Options passed to the extractor were being ignored by the rules, so configuration looked like it worked and didn’t. And the Replay button skipped to the end instead of replaying, which defeated the point of recording decisions in the first place.

What’s next

  • More adapters, starting with the ones people request
  • Better handling of the flaky-test and security questions, where trimming hurt

The code is on GitHub, MIT licensed, installable from PyPI, and the full documentation covers every adapter and the benchmark method. It is a community project, not affiliated with TypeSafe AI. See my other work, or read about Sable, where agent reliability is the whole point.

#agents#llm#python#open-source

Comments

All writing RSS Reply by email