<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://radar.montevive.ai/</id>
  <title>LLM Radar</title>
  <subtitle>What actually moved in LLMs, once a month. llama.cpp, small-model compression, and the open weights worth downloading.</subtitle>
  <link rel="self" type="application/atom+xml" href="https://radar.montevive.ai/feed.xml"/>
  <link rel="alternate" type="text/html" href="https://radar.montevive.ai/"/>
  <updated>2026-08-16T12:00:00Z</updated>
  <author>
    <name>Montevive</name>
    <uri>https://montevive.ai</uri>
    <email>ai@montevive.ai</email>
  </author>
  <rights>© 2026 Montevive</rights>

  <entry>
    <id>https://radar.montevive.ai/issues/2026-08-16-llm-state-of-the-art.html</id>
    <title>Issue 01 — Compression stopped being a research topic and became a runtime feature</title>
    <link rel="alternate" type="text/html" href="https://radar.montevive.ai/issues/2026-08-16-llm-state-of-the-art.html"/>
    <updated>2026-08-16T12:00:00Z</updated>
    <published>2026-08-16T12:00:00Z</published>
    <category term="llama.cpp"/>
    <category term="small language models"/>
    <category term="quantization"/>
    <summary type="text">Covering 16 July to 16 August 2026. llama.cpp got NVFP4 end to end plus a KV cache cloning endpoint reporting up to 2.12x on shared-prefix workloads, and suffix decoding brought speculative decoding without a draft model. On-policy distillation displaced structured pruning as the way to make a small model good, and KV cache compression is where the research volume went. Four independent papers found compressed models passing every standard quality guard while behaving measurably differently. Includes a glossary of the underlying technology.</summary>
  </entry>

</feed>
