DEV Community

#llm

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
llama.cpp vs Ollama — which one should you run?

llama.cpp vs Ollama — which one should you run?

5
Comments 1
5 min read
Why Byte-Faithful Pass-Through Beats Protocol Translation for Tool-Calling Agents

Why Byte-Faithful Pass-Through Beats Protocol Translation for Tool-Calling Agents

Comments 3
4 min read
MCP Connected Your Tools. It Didn't Fix Your Agent's Memory.

MCP Connected Your Tools. It Didn't Fix Your Agent's Memory.

3
Comments 2
4 min read
Benchmark Contamination 101: How Train/Test Overlap Inflates Leaderboard Scores (and How to Catch It)

Benchmark Contamination 101: How Train/Test Overlap Inflates Leaderboard Scores (and How to Catch It)

Comments
7 min read
KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly

KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly

Comments
7 min read
Restoring a grant is not restoring capacity

Restoring a grant is not restoring capacity

Comments
9 min read
When the attacker is an agent: a defender's field guide to autonomous AI intrusions in 2026

When the attacker is an agent: a defender's field guide to autonomous AI intrusions in 2026

Comments
6 min read
OpenAI Started Watermarking ChatGPT Text. Build a Tiny Text Watermark in TypeScript.

OpenAI Started Watermarking ChatGPT Text. Build a Tiny Text Watermark in TypeScript.

2
Comments 1
12 min read
Do LLMs Invent Japanese Law Articles? A Bilingual Benchmark

Do LLMs Invent Japanese Law Articles? A Bilingual Benchmark

Comments
4 min read
How to Build Resilient AI Agents with Search Fallback Loops

How to Build Resilient AI Agents with Search Fallback Loops

Comments
5 min read
4-bit GGUF Quality for MoE Models: Why Only 3B of 180B Params Fire, and How to Prove Parity

4-bit GGUF Quality for MoE Models: Why Only 3B of 180B Params Fire, and How to Prove Parity

Comments
7 min read
Indirect Prompt Injection Through Tool Descriptions and Tool Output: How Untrusted Metadata Hijacks Agents

Indirect Prompt Injection Through Tool Descriptions and Tool Output: How Untrusted Metadata Hijacks Agents

Comments
7 min read
Le Chonk Unleashed: Why Mistral’s 1‑Trillion‑Parameter Model Is Shattering LLM Benchmarks

Le Chonk Unleashed: Why Mistral’s 1‑Trillion‑Parameter Model Is Shattering LLM Benchmarks

Comments
5 min read
Claude Opus 5.5: What Changed, What Breaks, and Whether to Switch

Claude Opus 5.5: What Changed, What Breaks, and Whether to Switch

Comments
7 min read
ZTC (Zero-Token Confidence): juzgar una acción de IA leyendo el estado interno del modelo, con cero tokens extra

ZTC (Zero-Token Confidence): juzgar una acción de IA leyendo el estado interno del modelo, con cero tokens extra

Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.