jev-x-kitjev-x-kitView on GitHub ↗
v0.1.0 · MIT · 33 MCP tools

Your agent doesn't need
Opus to say yes.

jev-x-kit drops a non-autoregressive System-1 decision layer in front of every micro-decision your coding agent makes — Choice, Score and Noul primitives, schema-safe, $0 output tokens, 70–150ms. Runs 100% offline. No GPU. No API key required.

$ git clone https://github.com/Kadihx/jev-x-kit.git
View on GitHub ↗

~19x Fewer Tokens.
~11x Fewer Tool Calls.

*Measured on 3 real runs against live free APIs — not marketing copy. See the full methodology and reproduce every number yourself.

the coefficient

One float decides where every
decision goes.

Every call into jev-x-kit resolves to a single calibrated confidence score — the coefficient. RLCD (Reinforcement Learning for Calibrated Decisions) trains it to be statistically honest, so the number you get back is the number you can route on.

CHOICEprimitive
Up to 255 options

Picks from a defined option set. Returns probabilities + a calibrated confidence, never prose.

SCOREprimitive
Fractional 1–10

A sortable, comparable scalar (e.g. score: 7.235) instead of a paragraph of hedging.

NOULprimitive
Calibrated 0.0–1.0

A Bernoulli probability for yes/no gates — RLCD-trained to actually mean what it says.

BELKİ Gatekeeperrouting
  • confidence > 0.85execute directly · $0 LLM cost
  • 0.60 – 0.85speculative fan-out · N sub-decisions
  • < 0.60  (“BELKİ”)escalate · System 2 or human
Gatekeeper 1.1
confidence > 0.85 → execute
0.60 – 0.85 → speculative
< 0.60 → System 2
00:00:00
Backend chain coefficient$/1M tokens
OpenJev / vLLM100x$0 · Docker, local GPU
LayA (ModernBERT)100x$0 · ~5ms on Apple Silicon
Ollama (System-2)100x$0 · planner / red-team text
Vercel AI Gateway40xfree tier
TypeSafe Jev (native)1x$0.042 / 1M in · $0 out

Auto-resolution order: typesafe_jev → openjev_local → laya_local → heuristic. The offline heuristic simulator always terminates the chain, so nothing ever fails closed — even with no network, no GPU and no API key.

We took the boring
path on purpose.

Latency wall

3–30s for a single yes/no routing decision, every time, on the critical path.

Cost explosion

Agent loops burn $10–$50/session on micro-decisions routed to frontier models.

Schema breakage

Free-text models hallucinate parameters and quietly break JSON parsers.

Context rot

“Summarized” logs hallucinate file paths, commands and error codes.

Single-pass blind spot

A model can't adversarially review its own plan for race conditions.

Vendor lock-in

Paid-API-only tools leave offline and budget-constrained developers out.

Non-Autoregressive

Choice / Score / Noul — schema-validated, zero output tokens, zero parse errors.

Winnow, Not Summarize

Irrelevant log lines are deleted, never paraphrased. Paths, commands and error codes survive byte-identical.

Self-Improving

RLVR records tsc / test outcomes into .jev-skill-memory.json and auto-tunes the gatekeeper.

see it run

Real output. Same machine, zero network.

bash — jev info
$ node dist/cli.js info
{
  "backend": {
    "id": "typesafe_jev",
    "label": "TypeSafe Jev API (POST /v1/systemone)",
    "pricePerMillionUsd": 0.042,
    "outputTokenCostUsd": 0,
    "local": false
  },
  "chain": [
    "typesafe_jev @ https://api.typesafe.ai/v1",
    "openjev_local @ http://localhost:8000/v1",
    "laya_local @ http://localhost:8000/v1",
    "selected: typesafe_jev"
  ],
  "tools": 33
}
bash — jev decide
$ node dist/cli.js decide \
    "run tsc before pushing?" --options "yes,no"
{
  "decision": {
    "question": "run tsc before pushing?",
    "route": "system2",
    "confidence": 0.22,
    "belki": true,
    "reason": "confidence 0.22 < 0.6
      (\"BELKİ\") -> System 2 cascade"
  },
  "selected": "yes",
  "probabilities": [0.61, 0.39]
}

Screenshots of the MCP tool calls inside Claude Code go here — dropped in straight from real sessions, not mockups.

setup

Install as a Claude Code plugin

Thirty seconds. No account, no key required.

bash
git clone https://github.com/Kadihx/jev-x-kit.git
cd jev-x-kit && npm install && npm run build

npm test            # 43 unit tests, offline simulator
npm run smoke       # 69-check end-to-end MCP client test

.claude-plugin/plugin.json is already wired up — it registers the jev skill and the jev-super-agent MCP server with 33 tools. Point Claude Code, Cursor, OpenCode or Continue.dev at this folder as a plugin, or hand-wire the MCP config yourself:

mcp-config.json
{
  "mcpServers": {
    "jev-super-agent": {
      "command": "node",
      "args": ["/path/to/jev-x-kit/dist/index.js"],
      "env": {
        "JEV_BACKEND_PROVIDER": "auto",
        "OPENJEV_BASE_URL": "http://localhost:8000/v1",
        "JEV_LLM_BASE_URL": "http://localhost:11434/v1"
      }
    }
  }
}
one-click installnpm run install-mcpdetects Claude Desktop, Cursor and Continue.dev, merges (never overwrites) config, backs up originals to .bak.
usage tips

CLI cheat sheet

Scriptable, zero MCP client needed — every module is also a plain CLI command.

info
node dist/cli.js info

backend + chain diagnosis

decide
node dist/cli.js decide "run tsc first?" --options "yes,no"

BELKİ gatekeeper call

plan
node dist/cli.js plan "Ship an offline decision layer" --preset software-architecture

hypothesis + anti-thesis + Jev score

redteam
node dist/cli.js redteam "We cache everything forever"

adversarial dual loop

compact
node dist/cli.js compact --file build.log --goal "port binding error"

Winnow lossless compaction

audit
node dist/cli.js audit .

360° architecture / security / marketing scan

verify
node dist/cli.js verify --cwd .

RLVR: tsc + tests → reward

MCP tools — 33 total, highlights below

ToolModuleWhat it does
jev_evaluate1Fan-out batch of Choice/Score/Noul in one pass
jev_decide1BELKİ gatekeeper: execute / speculative / system2
jev_plan2Ultra-planning: hypothesis + anti-thesis + 4-dim score
jev_redteam3Adversarial dual loop: anti-theses, severity, arbitration
jev_audit4360° scan: architecture / security / marketing / legal / budget
jev_research54 parallel channels + Jev re-rank, ~19x fewer tokens
jev_compact5Winnow lossless context compaction (delete, never summarize)
jev_github_mine6License audit + clean-room originality guard
jev_verify8RLVR: run tsc/tests, record reward, auto-tune thresholds
jev_guardrail9AutoMode pre-execution gate (allow / ask / block)
hub_queryhubFTS5 + Jev-ranked, cited answers from the research hub
jev_competitor_scannewPaginated GitHub search, Noul-ranked against your own repo

7 multi-domain presets

software-architecturemarketing-growthproduct-uxcost-model-routercybersecuritylegal-compliancefinance-valuation

Every preset injects rules, red-flags, checklists and a scoring rubric straight into Jev state, verbatim — no prompt engineering required.

benchmark

jev_research vs. vanilla Claude Code

One jev_research call fans four channels out outside Claude's context, ranks them with a $0 local Noul primitive, and returns only the ranked shortlist plus a synthesized brief.

QueryCandidatesTokens (jev)Naive tokensReductionSpeed-up
rate limiting strategies for a public REST API122,41345,81719.0x3.5x
vector database comparison for RAG pipelines122,76445,81716.6x7.3x
how to reduce React app bundle size101,76438,18121.6x9.5x

Every number is either measured on real live free APIs, or derived from this repo's own crawled corpus — nowhere is a number invented. Wall-clock speed-up assumes 3s/manual turn (conservative, change it and re-run if you disagree). Reproduce any row yourself with node scripts/benchmark-vs-vanilla.mjs "your query". Full methodology in BENCHMARK.md.

benchmark

batch() vs. one-at-a-time: 4.24x

Same 10-question battery, same real typesafe_jev backend, two ways of calling it: naive sequential calls (what a raw API integration looks like without the kit) vs. jev-x-kit's own backend.batch() (same-state questions merge into one HTTP call, different-state questions run concurrently).

BackendSequential (10 calls)batch() (1 call)Speedup
heuristic (offline, $0)2.69ms0.36ms7.47x
typesafe_jev (real API)3,479ms821ms4.24x

LayA turns out to really exist — an open-source, self-hosted 421M ModernBERT + RLCD decision model (github.com/NandhaKishorM/laya) — so we installed it and ran the exact same 10-question battery against it, on this (GPU-less, CPU-only) machine:

Backendms / question (batched)correct (4 factual Qs)
typesafe_jev (real API)82.1ms4/4
laya (local, CPU)183.9ms2/4

Not a clean win for either side: this machine has no GPU (LayA publishes ~33–38ms/question on GPU, so these are CPU-bound numbers), and the battery is open-domain trivia — LayA's own model card lists its strengths as email triage, moderation and classification, not general knowledge, so this plays to a hosted LLM API's strengths more than LayA's. Reproduce it yourself with scripts/laya-benchmark.py; full numbers in artifacts/laya-benchmark-report.md.

Every number here is a real local measurement against the live TypeSafe Jev API, not a projection. Reproduce it yourself with node scripts/backend-benchmark.mjs (needs TYPESAFE_JEV_API_KEY in .env for the typesafe_jev row). Full methodology in artifacts/backend-benchmark-report.md.

Bonus: a research hub for rational thinking

A nightly-crawled personal-development and rationality corpus — 11 sources across mental-models, library and academic tiers, stored in SQLite + FTS5, ranked by Jev and cited VERIFIED / PROBABLE / REJECTED. Farnam Street, LessWrong, Derek Sivers, Julian Shapiro, Internet Archive, Gutenberg, Wikibooks, PhilArchive, PsyArXiv, CORE.

bash — hub
npm run hub -- crawl                          # nightly, 02:00–06:00 Berlin
npm run hub -- query "second-order thinking"  # FTS5 + Jev-ranked, cited

Give your agent a $0 System 1.

Star it on GitHub ↗
git clone github.com/Kadihx/jev-x-kit