Local inference 🧩 Products
Local inference is 💤 Quiet on Above the Fog’s AI hype radar: a hype score of 10 out of 100 on Sep 27, 2026, 8 points higher than 7 days before. Its attention is within the normal range across the monitored sources. Its last hype episode in the 30-day window ran from Sep 11 to Sep 11 (peak 29 on Sep 11).
Running models on your own machine or server: Ollama, llama.cpp, vLLM, MLX, LM Studio, and the GGUF format.
First seen on HN: front-page time (ClickHouse) · Sep 14 — the first monitored source to go above normal for it in the last 14 days.
Why this stage: within the normal range across monitored sources.
- Stage
- 💤 Quiet — within the normal range across the monitored sources
- Hype score
- 10 / 100 · −3.7 in 1 day · +8.2 in 7 days
- In hype since
- not in hype now
- Peak, 30 days
- 36.9 on Sep 4, 2026
- Category
- 🧩 Products
- Part of
- Open-weight models
- Sources above normal
- 1 of 27 scoring it (1 independent family)
- Stories in 7 days
- 646
- Share of voice
- 1.1% of the radar’s attention in 7 days
Where the attention comes from
“Normal” is the median of the source’s previous 14 normal days for this topic. A source is above normal when its value is at least 30% over that level and stands out from its usual day-to-day variation (z ≥ 2); the share is the part of the hype score that source explains. “× normal” is left blank when the normal level is 0.
| Source | Metric | Latest | Normal | × normal | Above normal | Share of score |
|---|---|---|---|---|---|---|
| HN: front-page time (ClickHouse) | frontpage hours | 6 | 0 | — | yes | 95% |
| DEV Community (dev.to) | items | 13 | 10.5 | 1.2× | no | 5.1% |
| Developer forums (Cursor, OpenAI, Hugging Face) | items | 1 | 0 | — | no | 0% |
| Luma (SF AI events) | events | 1 (Sep 26) | 0 | — | no | 0% |
| Bluesky | trending | 0 (Sep 26) | 0 | — | no | 0% |
| Cerebral Valley (AI events) | events | 0 (Sep 26) | 0 | — | no | 0% |
| GitHub Trending | trending | 0 (Sep 26) | 0 | — | no | 0% |
| Google Trends (Trending now) | trending | 0 (Sep 26) | 0 | — | no | 0% |
Stories behind it
- NVIDIA/Model-Optimizer: A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize — GitHub Trending (NVIDIA) · score 4,697
- 42x Faster Prompt Lookup Drafting in llama.cpp — Reddit (r/LocalLLaMA) · Sep 27 · score 607 · 172 comments
- Ollaya – Ollama for open-source, Jev-style decision models — Hacker News (ollaya.dev) · Sep 25 · score 609 · 145 comments
- ollaya-dev/ollaya: Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models. — GitHub (ollaya-dev) · Sep 23 · score 862
- Ollaya – Ollama for open-source, Jev-style decision models — HN: front-page time (ClickHouse) (ollaya.dev) · Sep 25 · score 568 · 137 comments
- OrcaSAQ-2-27B — Hugging Face Hub (trending + orgs) (orcarouter) · Sep 24 · score 187
- I am currently having access to an Apple Mac M4 Max with 36GB unified RAM and ohhhh gurl does it do local LLMs! :blobcatrainbow: It does Qwen3.8 27B Q4 at 18 tokens/s and only uses 40W... completely… — Mastodon (fediverse) (lgbtqia.space) · Sep 23 · score 3 · 2 comments
- Llama-modes: load one GGUF once, then Chat, Boolean, Choice and Scale — Developer forums (Cursor, OpenAI, Hugging Face) (Hugging Face Forums) · Sep 27 · score 1 · 1 comment
- Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching — Hugging Face Daily Papers (hf-papers) · Sep 28 · score 1 · 1 comment
- Watermarking in vLLM — Lobste.rs (#vibecoding) · Sep 24 · score 2 · 0 comments
- GPT-6 ⚡, Opus 5.5 🧠, AI leaders at UN 🌐 — AI newsletters (RSS) (TLDR AI) · Sep 23
- The Model Decision: Optimizing Cost, Control, & Speed — Luma (SF AI events) (Judy Wu) · score 0
Themes
What its stories of the last 7 days are about: the 553 of them the Jev model kept as about the scene, by theme (a story can carry several themes):
- Models — 84% of its stories
- Agents — 17% of its stories
- Protocols & infra — 17% of its stories
- Usage & adoption — 15% of its stories
- Security & safety — 2.9% of its stories
- Other tech — 0.9% of its stories
Related topics
- Related to
- Hermes Agent, Llama, Hugging Face, OpenCode
- Part of
- Open-weight models
- Mentioned with
- Qwen (514 stories), Gemma (94 stories), Decision models (32 stories), Mistral AI (20 stories), DeepSeek Harness (15 stories), Muse Spark (9 stories)
Hype score, last 30 days
| Day | Hype score | Stage |
|---|---|---|
| Sep 27, 2026 | 10 | 💤 Quiet |
| Sep 26, 2026 | 13.7 | 💤 Quiet |
| Sep 25, 2026 | 13 | 💤 Quiet |
| Sep 24, 2026 | 5.2 | 💤 Quiet |
| Sep 23, 2026 | 5.8 | 💤 Quiet |
| Sep 22, 2026 | 13.6 | 💤 Quiet |
| Sep 21, 2026 | 3.2 | 💤 Quiet |
| Sep 20, 2026 | 1.8 | 💤 Quiet |
| Sep 19, 2026 | 4.4 | 💤 Quiet |
| Sep 18, 2026 | 4.7 | 💤 Quiet |
| Sep 17, 2026 | 17.8 | 💤 Quiet |
| Sep 16, 2026 | 5.8 | 💤 Quiet |
| Sep 15, 2026 | 14.2 | 🧊 Cooling |
| Sep 14, 2026 | 17.6 | 🧊 Cooling |
| Sep 13, 2026 | 0 | 💤 Quiet |
| Sep 12, 2026 | 16.3 | 🧊 Cooling |
| Sep 11, 2026 | 28.7 | 🚀 Rising |
| Sep 10, 2026 | 22 | 🧊 Cooling |
| Sep 9, 2026 | 17.9 | 🧊 Cooling |
| Sep 8, 2026 | 16.5 | 🧊 Cooling |
| Sep 7, 2026 | 32.3 | 🔥 Hot |
| Sep 6, 2026 | 11.5 | 🧊 Cooling |
| Sep 5, 2026 | 21.6 | 🧊 Cooling |
| Sep 4, 2026 | 36.9 | 🚀 Rising |
| Sep 3, 2026 | 15.6 | 💤 Quiet |
| Sep 2, 2026 | 16.4 | 🧊 Cooling |
| Sep 1, 2026 | 16.8 | 🧊 Cooling |
| Aug 31, 2026 | 15.6 | 🧊 Cooling |
| Aug 30, 2026 | 11.8 | 🧊 Cooling |
| Aug 29, 2026 | 27.5 | 🔥 Hot |