DEV Community

#benchmarking

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Fan-Out: Concurrent Tool Calls Under Shared-State Hazards

Kaggle Benchmarking Challenge Submission

Fan-Out: Concurrent Tool Calls Under Shared-State Hazards

1
Comments 1
3 min read
Warmup Deep Dive: What You’re Throwing Away and Why

Warmup Deep Dive: What You’re Throwing Away and Why

Comments
4 min read
Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens

Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens

Comments
3 min read
Benchmarking regex patterns properly: P95/P99, equivalence checks, real engines

Benchmarking regex patterns properly: P95/P99, equivalence checks, real engines

Comments
2 min read
MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?

Kaggle Benchmarking Challenge Submission

MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?

Comments
3 min read
LiteLLM Rust Gateway Benchmarked: Fast and Tiny, but Not Yet a Python Proxy Replacement

LiteLLM Rust Gateway Benchmarked: Fast and Tiny, but Not Yet a Python Proxy Replacement

Comments
11 min read
I benchmarked what frontier models actually know about 2026 — most of it, they don't

Kaggle Benchmarking Challenge Submission

I benchmarked what frontier models actually know about 2026 — most of it, they don't

Comments
3 min read
Measuring the Rust Storage Engine: Lioran S3 PUT Timing from Network to RocksDB

Measuring the Rust Storage Engine: Lioran S3 PUT Timing from Network to RocksDB

1
Comments
2 min read
Kaniko's Performance Gap Narrows: Outdated 2018 Benchmarks Mislead, Recent Updates Close BuildKit Speed Divide

Kaniko's Performance Gap Narrows: Outdated 2018 Benchmarks Mislead, Recent Updates Close BuildKit Speed Divide

Comments
13 min read
Jev is the best decision model. Here's what to run when you can't use it.

Jev is the best decision model. Here's what to run when you can't use it.

Comments
10 min read
NeMo Guardrails vs Guardrails AI: The Production Latency Benchmark We Could Not Honestly Complete

NeMo Guardrails vs Guardrails AI: The Production Latency Benchmark We Could Not Honestly Complete

Comments
11 min read
I built a CoreMark leaderboard for 100+ NAS CPUs — here's what the data actually shows

I built a CoreMark leaderboard for 100+ NAS CPUs — here's what the data actually shows

Comments
2 min read
Warm vs Cold: Why a Single Trial Misleads Performance Claims

Warm vs Cold: Why a Single Trial Misleads Performance Claims

1
Comments
4 min read
When to use which: Dragonfly vs Redis vs Valkey

When to use which: Dragonfly vs Redis vs Valkey

1
Comments
4 min read
Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute — including ours

Our AI agents' "verified success" claims: 10 out of 10 failed independent recompute — including ours

1
Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.