Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
LLM-as-Judge Is Not a Score. It Is a Reasoning Contract.
Harrison Guo
Harrison Guo
Harrison Guo
Follow
Oct 6
LLM-as-Judge Is Not a Score. It Is a Reasoning Contract.
#
llm
#
ai
#
evaluation
#
calibration
Comments
Add Comment
5 min read
Sorry, That Was For Another Chat
JaviMaligno
JaviMaligno
JaviMaligno
Follow
Oct 6
Sorry, That Was For Another Chat
#
ai
#
agents
#
evaluation
#
claude
Comments
Add Comment
8 min read
Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances
orca forge
orca forge
orca forge
Follow
Oct 2
Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances
#
llm
#
ai
#
evaluation
#
japanese
Comments
Add Comment
12 min read
Done Is Testimony. Terminal State Is the Grade.
Igor Eduardo
Igor Eduardo
Igor Eduardo
Follow
Oct 5
Done Is Testimony. Terminal State Is the Grade.
#
ai
#
agents
#
evaluation
#
testing
1
 reaction
Comments
2
 comments
4 min read
Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements
Artyom Kornilov
Artyom Kornilov
Artyom Kornilov
Follow
Oct 1
Assessing the Security of a Nondeterministic Encryption Method: Identifying Weaknesses and Improvements
#
encryption
#
security
#
cryptography
#
evaluation
Comments
Add Comment
12 min read
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark
RESK
RESK
RESK
Follow
Oct 2
How LLM Evaluation Produces Comparable Numbers: Inside the LFORLA Reverse Engineering Benchmark
#
llm
#
evaluation
#
benchmark
#
reverseengineering
2
 reactions
Comments
Add Comment
4 min read
I rejected a model that passed everyone else's benchmark
Priyansh Kansara
Priyansh Kansara
Priyansh Kansara
Follow
Oct 3
I rejected a model that passed everyone else's benchmark
#
machinelearning
#
deepfake
#
evaluation
#
python
Comments
1
 comment
5 min read
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs
RESK
RESK
RESK
Follow
Sep 30
Election Forecasting Benchmarks: How LFORLA Scores Political Bias in LLMs
#
llm
#
benchmark
#
politics
#
evaluation
1
 reaction
Comments
Add Comment
4 min read
Your Agent Has Observability. It Doesn't Have Evals.
Jason Lau
Jason Lau
Jason Lau
Follow
Sep 24
Your Agent Has Observability. It Doesn't Have Evals.
#
agents
#
observability
#
evaluation
#
llm
Comments
Add Comment
10 min read
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges
Modusensus
Modusensus
Modusensus
Follow
Sep 25
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges
#
ai
#
oss
#
evaluation
#
llm
Comments
2
 comments
5 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
Amirul Cyber
Amirul Cyber
Amirul Cyber
Follow
Sep 17
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
#
ai
#
security
#
evaluation
#
llm
Comments
Add Comment
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
saaro
saaro
saaro
Follow
Sep 17
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
#
rag
#
ai
#
evaluation
#
metrics
Comments
Add Comment
4 min read
Using Execution Traces to Evaluate AI Agent Behavior
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 17
Using Execution Traces to Evaluate AI Agent Behavior
#
ai
#
opensource
#
evaluation
2
 reactions
Comments
Add Comment
5 min read
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)
Debashish Ghosal
Debashish Ghosal
Debashish Ghosal
Follow
Sep 24
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)
#
ai
#
evaluation
#
llm
#
programming
24
 reactions
Comments
5
 comments
5 min read
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
Cristian Gormaz
Cristian Gormaz
Cristian Gormaz
Follow
Sep 7
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
#
ai
#
testing
#
llm
#
evaluation
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account