DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Spec, Hash or Guess: Can LLMs Keep Spain's Tamper-Proof Invoice Ledger?

Kaggle Benchmarking Challenge Submission

Spec, Hash or Guess: Can LLMs Keep Spain's Tamper-Proof Invoice Ledger?

1
Comments
7 min read
Every LLM I tested would help my mother prepay a fake OLX seller

Kaggle Benchmarking Challenge Submission

Every LLM I tested would help my mother prepay a fake OLX seller

Comments
11 min read
Can an AI catch the catch? I benchmarked 14 models on bounty fine print

Kaggle Benchmarking Challenge Submission

Can an AI catch the catch? I benchmarked 14 models on bounty fine print

Comments
6 min read
Receipt Before Claim: a 4% score that did not mean 96% wrong reasoning

Kaggle Benchmarking Challenge Submission

Receipt Before Claim: a 4% score that did not mean 96% wrong reasoning

Comments
4 min read
I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.

Kaggle Benchmarking Challenge Submission

I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.

1
Comments
4 min read
The live realtor model passed. The goodbye failed. published: true

Kaggle Benchmarking Challenge Submission

The live realtor model passed. The goodbye failed. published: true

1
Comments
3 min read
My AI agents didn't fake citations. One in four still didn't hold.

Kaggle Benchmarking Challenge Submission

My AI agents didn't fake citations. One in four still didn't hold.

Comments
11 min read
Fan-Out: Concurrent Tool Calls Under Shared-State Hazards

Kaggle Benchmarking Challenge Submission

Fan-Out: Concurrent Tool Calls Under Shared-State Hazards

Comments
3 min read
Green Tests, Broken Architecture: Four Models Review TypeScript Diffs

Kaggle Benchmarking Challenge Submission

Green Tests, Broken Architecture: Four Models Review TypeScript Diffs

1
Comments
4 min read
Kaggle Benchmarking Challenge

Kaggle Benchmarking Challenge Submission

Kaggle Benchmarking Challenge

Comments
7 min read
TOOL JUDGMENT: Does an AI Know When NOT to Use a Tool?

TOOL JUDGMENT: Does an AI Know When NOT to Use a Tool?

Comments
5 min read
Asked in Dutch, sent to an English menu

Kaggle Benchmarking Challenge Submission

Asked in Dutch, sent to an English menu

2
Comments
8 min read
Alberta stopped changing its clocks in June. 19 of 19 frontier models still put Calgary on standard time in November.

Kaggle Benchmarking Challenge Submission

Alberta stopped changing its clocks in June. 19 of 19 frontier models still put Calgary on standard time in November.

5
Comments
15 min read
ContractClarity: Benchmarking LLMs on Contract Understanding

Kaggle Benchmarking Challenge Submission

ContractClarity: Benchmarking LLMs on Contract Understanding

Comments
3 min read
How well do AI models read Tanglish? I tested 13 of them

How well do AI models read Tanglish? I tested 13 of them

Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.