In one paragraph
A full LLM is slow and expensive when you only need one answer: is this PII, which option wins, how confident are we. Small decision models address that by specializing. pii runs a compiled neural detector that returns typed PII spans locally. strands-decider replaces the language-modelling head with a pointer head that scores options directly, giving calibrated confidence on every choice, score, or noul. awesome-jev catalogs public projects built on Jev, TypeSafe AI’s typed decision model. What counts here: small, fast models that return one typed answer or classification in place of a full LLM call.
The main approaches
Decision models
A model trained or fine-tuned to return a typed answer (choice, score, noul, confidence) rather than free text. The model reads a state once and answers typed questions against it. strands-labs/strands-decider is the example here: it uses a pointer head instead of a generation loop and reports 115 ms median latency per question on an RTX 3090.
Task-specific classifiers
A model compiled for one narrow task and shipped as a callable artifact. programasweights/pii fits here: one fixed compiled program, downloaded once, that returns typed PII spans (private_email, secret, account_number, etc.) with no configuration.
Ecosystems
A curated index tracking real-world usage of a specific decision model across many domains. yibie/awesome-jev catalogs projects, integrations, and discussions built on Jev, organized into 15 categories including Classification and Routing, Verification and Guardrails, and Agent Decisions.
Map of the theme
flowchart LR t["Small decision models"] t --> f1["Decision models"] t --> f2["Ecosystems"] t --> f3["Task-specific classifiers"] f1 --> r1["strands-labs/strands-decider"] f2 --> r2["yibie/awesome-jev"] f3 --> r3["programasweights/pii"]
Where the new ideas are
strands-labs/strands-decider replaces the language-modelling head with a pointer head that compares hidden states to score options. This means batching many typed questions against one text state costs little more than a single question, and every answer carries a calibrated confidence value without a generation loop.
programasweights/pii uses a fixed compiled program ID as the distribution unit. The model is not a weight file you configure; it is a named artifact you call. Benchmark numbers (0.8464 typed-character F1, 93.9% type accuracy) are published against the AI4Privacy dataset with disjoint splits for format selection and validation.
yibie/awesome-jev applies the ecosystem-list pattern to a single small model rather than a framework, requiring each entry to use Jev, or a documented port or derivative, for a concrete decision task.
Side by side
| Repo | Approach | Output types | Calibrated confidence | Local / self-hosted | Benchmark reported |
|---|---|---|---|---|---|
| programasweights/pii | Task-specific classifiers | 9 typed PII span labels | Not stated | Yes, one-time download | 0.8464 typed-char F1, 93.9% type accuracy (AI4Privacy) |
| yibie/awesome-jev | Ecosystems | Tracks Choice, Score, Boolean, Noul (via Jev) | Not stated | Not stated | Not stated |
| strands-labs/strands-decider | Decision models | Choice, Noul (0–1), Score | Yes, on every decision | Yes, runs on RTX 3090, Apple Silicon MLX, or CPU | 0.9+ confidence answers correct about 95% of the time on unseen short classification tasks; 115 ms median latency per question on RTX 3090 |
How the idea moved
flowchart LR n1["started Sep 2026<br/>programasweights/pii"] n2["started Sep 2026<br/>yibie/awesome-jev"] n3["started Sep 2026<br/>strands-labs/strands-decider"] n1 --> n2 --> n3
- started Sep 2026 · pii · Adds: A compiled, fixed-ID neural PII detector that runs locally, returns nine typed span labels, and reports typed-character F1 and type accuracy on the AI4Privacy dataset.
- started Sep 2026 · awesome-jev · Adds: A curated index, with its README built from per-category files, of 15 application categories showing how Jev’s typed decision model (Choice, Score, Boolean, Noul) is used in real projects and integrations.
- started Sep 2026 · strands-decider · Adds: A decision model with a pointer head instead of a generation loop that batches choice, score, and noul questions against one text state and returns a calibrated confidence value on every answer.
Easily confused
This theme might be confused with prompt-engineering guides or LLM fine-tuning tutorials, both of which also try to get structured answers from language models. The difference is that small decision models are separate, smaller artifacts specialized for typed answers rather than free text. They do not wrap a general-purpose LLM; they replace the LLM call for that decision. The theme is also distinct from agent frameworks: an agent framework orchestrates steps; a decision model is one fast step inside that orchestration.
Gaps nobody has filled
- None of these repos reports calibration separately for choice, score and noul answers. JevBench scores calibration for decision models but sits outside this theme.
- No repo handles streaming or incremental decisions where state changes between questions.
- No repo covers training a new task-specific compiled program from scratch without a vendor account; programasweights/pii supports recompiling from a plain-English spec but requires a PAW account and API key.
- None of these repos reports latency or accuracy on mobile or embedded hardware, though awesome-jev links to Android and ESP32 projects.