jev-x-kitView on GitHub ↗Your agent doesn't need
Opus to say yes.
jev-x-kit drops a non-autoregressive System-1 decision layer in front of every micro-decision your coding agent makes — Choice, Score and Noul primitives, schema-safe, $0 output tokens, 70–150ms. Runs 100% offline. No GPU. No API key required.
~19x Fewer Tokens.
~11x Fewer Tool Calls.
*Measured on 3 real runs against live free APIs — not marketing copy. See the full methodology and reproduce every number yourself.
One float decides where every
decision goes.
Every call into jev-x-kit resolves to a single calibrated confidence score — the coefficient. RLCD (Reinforcement Learning for Calibrated Decisions) trains it to be statistically honest, so the number you get back is the number you can route on.
Picks from a defined option set. Returns probabilities + a calibrated confidence, never prose.
A sortable, comparable scalar (e.g. score: 7.235) instead of a paragraph of hedging.
A Bernoulli probability for yes/no gates — RLCD-trained to actually mean what it says.
- confidence > 0.85execute directly · $0 LLM cost
- 0.60 – 0.85speculative fan-out · N sub-decisions
- < 0.60 (“BELKİ”)escalate · System 2 or human
Auto-resolution order: typesafe_jev → openjev_local → laya_local → heuristic. The offline heuristic simulator always terminates the chain, so nothing ever fails closed — even with no network, no GPU and no API key.
We took the boring
path on purpose.
3–30s for a single yes/no routing decision, every time, on the critical path.
Agent loops burn $10–$50/session on micro-decisions routed to frontier models.
Free-text models hallucinate parameters and quietly break JSON parsers.
“Summarized” logs hallucinate file paths, commands and error codes.
A model can't adversarially review its own plan for race conditions.
Paid-API-only tools leave offline and budget-constrained developers out.
Choice / Score / Noul — schema-validated, zero output tokens, zero parse errors.
Irrelevant log lines are deleted, never paraphrased. Paths, commands and error codes survive byte-identical.
RLVR records tsc / test outcomes into .jev-skill-memory.json and auto-tunes the gatekeeper.
Real output. Same machine, zero network.
$ node dist/cli.js info
{
"backend": {
"id": "typesafe_jev",
"label": "TypeSafe Jev API (POST /v1/systemone)",
"pricePerMillionUsd": 0.042,
"outputTokenCostUsd": 0,
"local": false
},
"chain": [
"typesafe_jev @ https://api.typesafe.ai/v1",
"openjev_local @ http://localhost:8000/v1",
"laya_local @ http://localhost:8000/v1",
"selected: typesafe_jev"
],
"tools": 33
}$ node dist/cli.js decide \
"run tsc before pushing?" --options "yes,no"
{
"decision": {
"question": "run tsc before pushing?",
"route": "system2",
"confidence": 0.22,
"belki": true,
"reason": "confidence 0.22 < 0.6
(\"BELKİ\") -> System 2 cascade"
},
"selected": "yes",
"probabilities": [0.61, 0.39]
}Screenshots of the MCP tool calls inside Claude Code go here — dropped in straight from real sessions, not mockups.
Install as a Claude Code plugin
Thirty seconds. No account, no key required.
git clone https://github.com/Kadihx/jev-x-kit.git cd jev-x-kit && npm install && npm run build npm test # 43 unit tests, offline simulator npm run smoke # 69-check end-to-end MCP client test
.claude-plugin/plugin.json is already wired up — it registers the jev skill and the jev-super-agent MCP server with 33 tools. Point Claude Code, Cursor, OpenCode or Continue.dev at this folder as a plugin, or hand-wire the MCP config yourself:
{
"mcpServers": {
"jev-super-agent": {
"command": "node",
"args": ["/path/to/jev-x-kit/dist/index.js"],
"env": {
"JEV_BACKEND_PROVIDER": "auto",
"OPENJEV_BASE_URL": "http://localhost:8000/v1",
"JEV_LLM_BASE_URL": "http://localhost:11434/v1"
}
}
}
}npm run install-mcpdetects Claude Desktop, Cursor and Continue.dev, merges (never overwrites) config, backs up originals to .bak.CLI cheat sheet
Scriptable, zero MCP client needed — every module is also a plain CLI command.
node dist/cli.js infobackend + chain diagnosis
node dist/cli.js decide "run tsc first?" --options "yes,no"BELKİ gatekeeper call
node dist/cli.js plan "Ship an offline decision layer" --preset software-architecturehypothesis + anti-thesis + Jev score
node dist/cli.js redteam "We cache everything forever"adversarial dual loop
node dist/cli.js compact --file build.log --goal "port binding error"Winnow lossless compaction
node dist/cli.js audit .360° architecture / security / marketing scan
node dist/cli.js verify --cwd .RLVR: tsc + tests → reward
MCP tools — 33 total, highlights below
| Tool | Module | What it does |
|---|---|---|
| jev_evaluate | 1 | Fan-out batch of Choice/Score/Noul in one pass |
| jev_decide | 1 | BELKİ gatekeeper: execute / speculative / system2 |
| jev_plan | 2 | Ultra-planning: hypothesis + anti-thesis + 4-dim score |
| jev_redteam | 3 | Adversarial dual loop: anti-theses, severity, arbitration |
| jev_audit | 4 | 360° scan: architecture / security / marketing / legal / budget |
| jev_research | 5 | 4 parallel channels + Jev re-rank, ~19x fewer tokens |
| jev_compact | 5 | Winnow lossless context compaction (delete, never summarize) |
| jev_github_mine | 6 | License audit + clean-room originality guard |
| jev_verify | 8 | RLVR: run tsc/tests, record reward, auto-tune thresholds |
| jev_guardrail | 9 | AutoMode pre-execution gate (allow / ask / block) |
| hub_query | hub | FTS5 + Jev-ranked, cited answers from the research hub |
| jev_competitor_scan | new | Paginated GitHub search, Noul-ranked against your own repo |
7 multi-domain presets
Every preset injects rules, red-flags, checklists and a scoring rubric straight into Jev state, verbatim — no prompt engineering required.
jev_research vs. vanilla Claude Code
One jev_research call fans four channels out outside Claude's context, ranks them with a $0 local Noul primitive, and returns only the ranked shortlist plus a synthesized brief.
| Query | Candidates | Tokens (jev) | Naive tokens | Reduction | Speed-up |
|---|---|---|---|---|---|
| rate limiting strategies for a public REST API | 12 | 2,413 | 45,817 | 19.0x | 3.5x |
| vector database comparison for RAG pipelines | 12 | 2,764 | 45,817 | 16.6x | 7.3x |
| how to reduce React app bundle size | 10 | 1,764 | 38,181 | 21.6x | 9.5x |
Every number is either measured on real live free APIs, or derived from this repo's own crawled corpus — nowhere is a number invented. Wall-clock speed-up assumes 3s/manual turn (conservative, change it and re-run if you disagree). Reproduce any row yourself with node scripts/benchmark-vs-vanilla.mjs "your query". Full methodology in BENCHMARK.md.
batch() vs. one-at-a-time: 4.24x
Same 10-question battery, same real typesafe_jev backend, two ways of calling it: naive sequential calls (what a raw API integration looks like without the kit) vs. jev-x-kit's own backend.batch() (same-state questions merge into one HTTP call, different-state questions run concurrently).
| Backend | Sequential (10 calls) | batch() (1 call) | Speedup |
|---|---|---|---|
| heuristic (offline, $0) | 2.69ms | 0.36ms | 7.47x |
| typesafe_jev (real API) | 3,479ms | 821ms | 4.24x |
LayA turns out to really exist — an open-source, self-hosted 421M ModernBERT + RLCD decision model (github.com/NandhaKishorM/laya) — so we installed it and ran the exact same 10-question battery against it, on this (GPU-less, CPU-only) machine:
| Backend | ms / question (batched) | correct (4 factual Qs) |
|---|---|---|
| typesafe_jev (real API) | 82.1ms | 4/4 |
| laya (local, CPU) | 183.9ms | 2/4 |
Not a clean win for either side: this machine has no GPU (LayA publishes ~33–38ms/question on GPU, so these are CPU-bound numbers), and the battery is open-domain trivia — LayA's own model card lists its strengths as email triage, moderation and classification, not general knowledge, so this plays to a hosted LLM API's strengths more than LayA's. Reproduce it yourself with scripts/laya-benchmark.py; full numbers in artifacts/laya-benchmark-report.md.
Every number here is a real local measurement against the live TypeSafe Jev API, not a projection. Reproduce it yourself with node scripts/backend-benchmark.mjs (needs TYPESAFE_JEV_API_KEY in .env for the typesafe_jev row). Full methodology in artifacts/backend-benchmark-report.md.
Bonus: a research hub for rational thinking
A nightly-crawled personal-development and rationality corpus — 11 sources across mental-models, library and academic tiers, stored in SQLite + FTS5, ranked by Jev and cited VERIFIED / PROBABLE / REJECTED. Farnam Street, LessWrong, Derek Sivers, Julian Shapiro, Internet Archive, Gutenberg, Wikibooks, PhilArchive, PsyArXiv, CORE.
npm run hub -- crawl # nightly, 02:00–06:00 Berlin npm run hub -- query "second-order thinking" # FTS5 + Jev-ranked, cited