Above the Fog

Local inference 🧩 Products

Local inference is 💤 Quiet on Above the Fog’s AI hype radar: a hype score of 10 out of 100 on Sep 27, 2026, 8 points higher than 7 days before. Its attention is within the normal range across the monitored sources. Its last hype episode in the 30-day window ran from Sep 11 to Sep 11 (peak 29 on Sep 11).

Running models on your own machine or server: Ollama, llama.cpp, vLLM, MLX, LM Studio, and the GGUF format.

First seen on HN: front-page time (ClickHouse) · Sep 14 — the first monitored source to go above normal for it in the last 14 days.

Why this stage: within the normal range across monitored sources.

Stage
💤 Quiet — within the normal range across the monitored sources
Hype score
10 / 100 · −3.7 in 1 day · +8.2 in 7 days
In hype since
not in hype now
Peak, 30 days
36.9 on Sep 4, 2026
Category
🧩 Products
Part of
Open-weight models
Sources above normal
1 of 27 scoring it (1 independent family)
Stories in 7 days
646
Share of voice
1.1% of the radar’s attention in 7 days

Where the attention comes from

“Normal” is the median of the source’s previous 14 normal days for this topic. A source is above normal when its value is at least 30% over that level and stands out from its usual day-to-day variation (z ≥ 2); the share is the part of the hype score that source explains. “× normal” is left blank when the normal level is 0.

SourceMetricLatestNormal× normalAbove normalShare of score
HN: front-page time (ClickHouse)frontpage hours60—yes95%
DEV Community (dev.to)items1310.51.2×no5.1%
Developer forums (Cursor, OpenAI, Hugging Face)items10—no0%
Luma (SF AI events)events1 (Sep 26)0—no0%
Blueskytrending0 (Sep 26)0—no0%
Cerebral Valley (AI events)events0 (Sep 26)0—no0%
GitHub Trendingtrending0 (Sep 26)0—no0%
Google Trends (Trending now)trending0 (Sep 26)0—no0%

Stories behind it

  1. NVIDIA/Model-Optimizer: A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize — GitHub Trending (NVIDIA) · score 4,697
  2. 42x Faster Prompt Lookup Drafting in llama.cpp — Reddit (r/LocalLLaMA) · Sep 27 · score 607 · 172 comments
  3. Ollaya – Ollama for open-source, Jev-style decision models — Hacker News (ollaya.dev) · Sep 25 · score 609 · 145 comments
  4. ollaya-dev/ollaya: Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models. — GitHub (ollaya-dev) · Sep 23 · score 862
  5. Ollaya – Ollama for open-source, Jev-style decision models — HN: front-page time (ClickHouse) (ollaya.dev) · Sep 25 · score 568 · 137 comments
  6. OrcaSAQ-2-27B — Hugging Face Hub (trending + orgs) (orcarouter) · Sep 24 · score 187
  7. I am currently having access to an Apple Mac M4 Max with 36GB unified RAM and ohhhh gurl does it do local LLMs! :blobcatrainbow: It does Qwen3.8 27B Q4 at 18 tokens/s and only uses 40W... completely… — Mastodon (fediverse) (lgbtqia.space) · Sep 23 · score 3 · 2 comments
  8. Llama-modes: load one GGUF once, then Chat, Boolean, Choice and Scale — Developer forums (Cursor, OpenAI, Hugging Face) (Hugging Face Forums) · Sep 27 · score 1 · 1 comment
  9. Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching — Hugging Face Daily Papers (hf-papers) · Sep 28 · score 1 · 1 comment
  10. Watermarking in vLLM — Lobste.rs (#vibecoding) · Sep 24 · score 2 · 0 comments
  11. GPT-6 ⚡, Opus 5.5 🧠, AI leaders at UN 🌐 — AI newsletters (RSS) (TLDR AI) · Sep 23
  12. The Model Decision: Optimizing Cost, Control, & Speed — Luma (SF AI events) (Judy Wu) · score 0

Themes

What its stories of the last 7 days are about: the 553 of them the Jev model kept as about the scene, by theme (a story can carry several themes):

Related topics

Related to
Hermes Agent, Llama, Hugging Face, OpenCode
Part of
Open-weight models
Mentioned with
Qwen (514 stories), Gemma (94 stories), Decision models (32 stories), Mistral AI (20 stories), DeepSeek Harness (15 stories), Muse Spark (9 stories)

Hype score, last 30 days

DayHype scoreStage
Sep 27, 202610💤 Quiet
Sep 26, 202613.7💤 Quiet
Sep 25, 202613💤 Quiet
Sep 24, 20265.2💤 Quiet
Sep 23, 20265.8💤 Quiet
Sep 22, 202613.6💤 Quiet
Sep 21, 20263.2💤 Quiet
Sep 20, 20261.8💤 Quiet
Sep 19, 20264.4💤 Quiet
Sep 18, 20264.7💤 Quiet
Sep 17, 202617.8💤 Quiet
Sep 16, 20265.8💤 Quiet
Sep 15, 202614.2🧊 Cooling
Sep 14, 202617.6🧊 Cooling
Sep 13, 20260💤 Quiet
Sep 12, 202616.3🧊 Cooling
Sep 11, 202628.7🚀 Rising
Sep 10, 202622🧊 Cooling
Sep 9, 202617.9🧊 Cooling
Sep 8, 202616.5🧊 Cooling
Sep 7, 202632.3🔥 Hot
Sep 6, 202611.5🧊 Cooling
Sep 5, 202621.6🧊 Cooling
Sep 4, 202636.9🚀 Rising
Sep 3, 202615.6💤 Quiet
Sep 2, 202616.4🧊 Cooling
Sep 1, 202616.8🧊 Cooling
Aug 31, 202615.6🧊 Cooling
Aug 30, 202611.8🧊 Cooling
Aug 29, 202627.5🔥 Hot

Open Local inference in the interactive map →