QUICK ANSWER

The models in AI Search’s September roundup are not interchangeable. Jev is a hosted, typed decision service billed on input tokens; Needle 3 is a tiny on-device tool, extraction and embedding model; MiniCPM5-2B is a general text generator you can run locally with a suitable runtime. Pick by task, then test quality and total cost on your own workload.

The small-model story is really three different stories

AI Search’s September 20 roundup placed Jev and Needle 3 alongside a wave of new models. Hugging Face’s live trending page is a useful discovery signal, but rank and download counts change; popularity does not prove accuracy or fit. The more useful pattern is that “small AI” now covers very different products: hosted decision APIs, tiny embedded specialists, and compact generative models. Compare them by what they do, what evidence supports them, and what your actual workload costs.

Needle 3: tiny specialist for local structured work

Cactus describes Needle 3 as an on-device model for tool selection, structured extraction, classification and text embeddings. Its model card reports a single 8–29 MB model file and a 121M-parameter full model; Cactus says the architecture’s effective arithmetic is closer to 50M parameters. It is not a general-purpose chat replacement. The intended benefit is handling narrow, schema-shaped jobs on phones, wearables, robots and other constrained devices without sending every request to a cloud API.

  • Hardware: designed for device-side runtimes; the model card does not prescribe a discrete GPU. CPU, mobile accelerator or supported embedded runtime may fit, depending on platform and latency target.
  • Price: no per-token hosted inference price is listed on the model card. Local inference avoids a provider token bill, while still using device power and engineering time.
  • Evidence: the reported exact-match tool-call and field-level extraction results are Cactus’s own full-split evaluations, not an independent FyreLinkz benchmark.

Jev: pay-per-input-token decision model, not a tiny local LLM

TypeSafe documents Jev 1.13 as a hosted System One model that returns calibrated, structured decisions. It accepts text, not image, audio or video input. Current listed price is $0.042 per million input tokens; output tokens are free. This makes cost estimation straightforward: 100 million input tokens cost $4.20; one billion cost $42, before any applicable tax or account-specific terms. For one million requests averaging 1,000 input tokens, that is also one billion input tokens, or $42 at the published rate. Token batching and repeated state affect the count, so use actual request payloads to estimate.

  • Use it when a workflow needs typed choices or rubric-style classification and you have checked its decisions on representative examples.
  • Do not compare its price directly with local generation: Jev sells a hosted decision service; a local model has hardware, energy and maintenance costs instead of a provider token charge.
  • Pin a version for evaluations. TypeSafe says aliases can move as releases change.

MiniCPM5-2B: a local generator with a small-model footprint

OpenBMB positions MiniCPM5-2B as a dense 2B text model for local and resource-constrained use, with tool-use and long-context goals. In OpenBMB’s own comparison table, it reports a 53.9 average against 33.2 for LiquidAI LFM2.5-2.6B and 28.0 for Qwen3.5-2B; on LiveCodeBench v6, it reports 69.1 against 42.1 for LFM2.5. These are the model maker’s selected evaluation results, not an independent or universal ranking, and it does not lead every individual test. Hugging Face lists Apache-2.0 for the model. A GGUF build lets compatible llama.cpp, Ollama or LM Studio setups load a quantized version; the repository documents a Q4_K_M quick start.

  • GPU: no single GPU is required if you run on CPU, but GPU offload can improve speed when your runtime supports your card. Start with the quantized file size, then reserve extra RAM or VRAM for runtime, context/KV cache and other processes.
  • Cost: local runs have no per-token provider fee. Estimate power as device watts × hours ÷ 1,000 × your electricity price per kWh; include the hardware purchase only if you are comparing total ownership cost.
  • Benchmark honestly: test your own prompts, language, context length and tool calls. Measure quality, peak memory and tokens per second on your hardware.

Benchmarks need a reliability and task-fit check

Two September arXiv preprints add a useful caution to Jev’s “typed” output story. One reports that, in its controlled rubric experiment, changing semantically loaded option names affected decisions even while the output remained type-valid. A second study comparing Jev with several LLM judges reports mixed outcomes across rubric panels, rather than one winner for every grading task. These are new preprints, not settled consensus and not proof that Jev fails in every use. They do show why schema-valid output must not be mistaken for a correct judgment. Use neutral option labels, blind spot checks and a human escalation path for consequential decisions.

Which one should you try first?

Use this shortlist to choose a first experiment; it is FyreLinkz editorial guidance, not a claim that we independently benchmarked the models.

NeedStarting pointHardware or spendWhat to verify
Route local app tools; extract fields on a constrained deviceNeedle 3On-device runtime; no published token rateExact-match tool calls, field F1, confidence calibration
Classify or score text with a typed decision outputJev 1.13Hosted API; $0.042 per million input tokens, output freeAccuracy by class, option-name bias, current token volume
Generate and edit text locally; explore tool useMiniCPM5-2B GGUFCPU or supported GPU; quantized model plus runtime/context memoryTask quality, peak memory, speed, license and deployment fit
Compare another compact local agent modelLiquidAI LFM2.5-2.6BLocal device; vendor reports under 2.5 GB memory for its setupReproduce results on your own device and prompts

Build a small evaluation you can reproduce

Create a small private test set from the task you actually plan to ship. Give every candidate the same inputs, freeze model versions and record answers before looking at aggregate scores. For each local run, note device, runtime, quantization, context length, peak memory and median/p95 latency. For Jev, count input tokens and apply the published rate; check the current docs before budgeting because prices and rate limits can change. Keep failures in the report. A model that costs less but causes expensive review or correction may not be cheaper overall.

REPRODUCIBLE CHECKLIST
  • Write 30–100 representative examples, including ambiguous and out-of-scope cases.
  • Define success before running: exact match, field-level F1, grounded answer, or human-rated quality.
  • Log version, prompt, runtime, quantization, hardware, latency and memory for each run.
  • Calculate cloud input-token cost and local energy separately; include human correction time.
  • Check confidence calibration and route uncertain or high-impact results to a person.

What this week’s trend means for buyers and builders

The model-count race is less useful than the task-fit shift. Needle 3 shows why a model can be tiny by specializing; Jev shows a hosted decision product can price input tokens rather than output prose; MiniCPM5-2B and LFM2.5 show continuing investment in capable local agents. None removes evaluation work. Choose the narrowest system that passes your test set, then compare its whole operating cost. For local AI planning, use our AI workstation planner to shortlist parts by country, budget and workload; for creators choosing hosted media tools, see our AI video model comparison.

FOLLOW THE SOURCE

Sources & further reading

Primary sources checked Sep 26, 2026. Vendor statements are attributed; editorial advice is our own.

  1. 1
    Weekly AI roundup, September 20, 2026 ↗AI Search on X · Sep 20, 2026
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
    JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places ↗Delip Rao and co-author · arXiv preprint · Sep 24, 2026
  10. 10
    Deploy local agents everywhere with LFM2.5-2.6B ↗Liquid AI · Hugging Face · Aug 4, 2026