DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
LLM-as-Judge Is Not a Score. It Is a Reasoning Contract.

LLM-as-Judge Is Not a Score. It Is a Reasoning Contract.

Comments
5 min read
Sorry, That Was For Another Chat

Sorry, That Was For Another Chat

Comments
8 min read
Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances

Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances

Comments
12 min read
Done Is Testimony. Terminal State Is the Grade.

Done Is Testimony. Terminal State Is the Grade.

1
Comments 2
4 min read
Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements

Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements

Comments
12 min read
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark

2
Comments
4 min read
I rejected a model that passed everyone else's benchmark

I rejected a model that passed everyone else's benchmark

Comments 1
5 min read
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs

1
Comments
4 min read
Your Agent Has Observability. It Doesn't Have Evals.

Your Agent Has Observability. It Doesn't Have Evals.

Comments
10 min read
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Comments 2
5 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

Comments
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

Comments
4 min read
Using Execution Traces to Evaluate AI Agent Behavior

Using Execution Traces to Evaluate AI Agent Behavior

2
Comments
5 min read
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

24
Comments 5
5 min read
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

Comments
4 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.