LLM Radar

What actually moved in LLMs, once a month — biased toward what you can actually run. llama.cpp, small-model compression, and the open weights worth the download.

Published by Montevive · Monthly · Free to read

What this is

What it covers

A fixed one-month window, read narrowly: what landed in llama.cpp, what the compression and small-model literature actually established, and which open weights shipped. Plus a glossary of the underlying technology.

Who it is for

People who deploy this stuff rather than only read about it — anyone choosing a quantization format, sizing a local model, or deciding whether a technique is ready. It assumes you build.

Who writes it

Montevive, an AI engineering studio in Granada. We build and run LLM systems for clients, and this is the digest we wanted to read and could not find.

How it is made

Most AI roundups are a model summarising other summaries. This one has a method, and the method is the point — it is why the numbers in it can be trusted.

The window is fixed before the research starts, and printed on the issue. Anything older that earns a mention is labelled as context, never smuggled in as news.

Every claim carries a link. No link, no claim. Papers cite the arXiv ID, llama.cpp work cites the pull request, releases cite the model card.

Load-bearing claims are verified at the primary source, not at a summary of it. In issue 01 that caught a widely-reported pull request described as shipped when it had in fact been rejected, and two numbers that search summaries had paraphrased wrongly.

Authors' own figures are labelled as claims, not measurements. A reported 20× compression is what the paper reports, and the issue says so.

Judgements are explicit and rankable — every recommendation carries an ADOPT, WATCH or LAB ONLY verdict, and the issue is willing to say lab only.

Current issue

Issue 0116 Jul — 16 Aug 20266 sections · ~70 citations

Compression stopped being a research topic and became a runtime feature.

  • llama.cpp got NVFP4 end to end — and a KV cache cloning endpoint that reports up to 2.12× on shared-prefix workloads.
  • Speculative decoding without a draft model, via suffix trees.
  • On-policy distillation displaced "prune the big one" as the way to make a small model good.
  • Four independent papers found compressed models passing every standard quality guard while behaving measurably differently.
Read issue 01

All issues

Get the next one

There is no signup form yet — and rather than collect addresses without a proper consent and unsubscribe path, we would rather you just ask. Email us to be added, or subscribe to the Atom feed, which needs no address at all.