DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Which AI lies less in marketing copy?

kagglechallenge

Which AI lies less in marketing copy?

1
Comments
16 min read
Every LLM I tested would help my mother prepay a fake OLX seller

Kaggle Benchmarking Challenge Submission

Every LLM I tested would help my mother prepay a fake OLX seller

Comments
11 min read
Can an AI catch the catch? I benchmarked 14 models on bounty fine print

Kaggle Benchmarking Challenge Submission

Can an AI catch the catch? I benchmarked 14 models on bounty fine print

Comments
6 min read
Receipt Before Claim: a 4% score that did not mean 96% wrong reasoning

Kaggle Benchmarking Challenge Submission

Receipt Before Claim: a 4% score that did not mean 96% wrong reasoning

Comments
4 min read
I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.

Kaggle Benchmarking Challenge Submission

I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.

1
Comments
4 min read
The live realtor model passed. The goodbye failed. published: true

Kaggle Benchmarking Challenge Submission

The live realtor model passed. The goodbye failed. published: true

1
Comments
3 min read
My AI agents didn't fake citations. One in four still didn't hold.

Kaggle Benchmarking Challenge Submission

My AI agents didn't fake citations. One in four still didn't hold.

Comments
11 min read
Kaggle Benchmarking Challenge

Kaggle Benchmarking Challenge Submission

Kaggle Benchmarking Challenge

Comments
7 min read
Asked in Dutch, sent to an English menu

Kaggle Benchmarking Challenge Submission

Asked in Dutch, sent to an English menu

2
Comments
8 min read
Alberta stopped changing its clocks in June. 19 of 19 frontier models still put Calgary on standard time in November.

Kaggle Benchmarking Challenge Submission

Alberta stopped changing its clocks in June. 19 of 19 frontier models still put Calgary on standard time in November.

5
Comments
15 min read
ContractClarity: Benchmarking LLMs on Contract Understanding

Kaggle Benchmarking Challenge Submission

ContractClarity: Benchmarking LLMs on Contract Understanding

Comments
3 min read
How well do AI models read Tanglish? I tested 13 of them

How well do AI models read Tanglish? I tested 13 of them

Comments
5 min read
FairHire-ES: Maternity-gap lines move Spanish résumé scores more than names (256 twin cases)

Kaggle Benchmarking Challenge Submission

FairHire-ES: Maternity-gap lines move Spanish résumé scores more than names (256 twin cases)

Comments
5 min read
Day 5: Professor Oakenscroll or: How I Hashed Every Line on My Laptop and Cut It Down to One Stranger's Photo, and... It Doesn't Matter

Day 5: Professor Oakenscroll or: How I Hashed Every Line on My Laptop and Cut It Down to One Stranger's Photo, and... It Doesn't Matter

Comments
6 min read
Would an AI let my patient drink thin water? Benchmarking LLMs on dysphagia diet safety

Kaggle Benchmarking Challenge Submission

Would an AI let my patient drink thin water? Benchmarking LLMs on dysphagia diet safety

Comments
13 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.