DEV Community

Machine Learning

A branch of artificial intelligence (AI) and computer science which focuses on the use of data and algorithms to imitate the way that humans learn, gradually improving its accuracy.

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Spec, Hash or Guess: Can LLMs Keep Spain's Tamper-Proof Invoice Ledger?

Kaggle Benchmarking Challenge Submission

Spec, Hash or Guess: Can LLMs Keep Spain's Tamper-Proof Invoice Ledger?

1
Comments
7 min read
Every LLM I tested would help my mother prepay a fake OLX seller

Kaggle Benchmarking Challenge Submission

Every LLM I tested would help my mother prepay a fake OLX seller

Comments
11 min read
I Diffed 4 CLAUDE.md Files From 4 Real Projects. Byte Range: 15 to 14,540.

I Diffed 4 CLAUDE.md Files From 4 Real Projects. Byte Range: 15 to 14,540.

Comments
4 min read
Can an AI catch the catch? I benchmarked 14 models on bounty fine print

Kaggle Benchmarking Challenge Submission

Can an AI catch the catch? I benchmarked 14 models on bounty fine print

Comments
6 min read
Stop Asking Your AI to Be Helpful. Ask It to Be Accountable.

Stop Asking Your AI to Be Helpful. Ask It to Be Accountable.

Comments
3 min read
I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.

Kaggle Benchmarking Challenge Submission

I tested 11 AI models on Indian GST, UPI and lakh-crore. Three famous ones got Puducherry wrong.

1
Comments
4 min read
Can You Gaslight an AI About AWS? I Measured It

Can You Gaslight an AI About AWS? I Measured It

Comments
8 min read
The World's Most Analyzed Dataset (and the Custody Chain Nobody Maintained)

The World's Most Analyzed Dataset (and the Custody Chain Nobody Maintained)

Comments
3 min read
The live realtor model passed. The goodbye failed. published: true

Kaggle Benchmarking Challenge Submission

The live realtor model passed. The goodbye failed. published: true

1
Comments
3 min read
My AI agents didn't fake citations. One in four still didn't hold.

Kaggle Benchmarking Challenge Submission

My AI agents didn't fake citations. One in four still didn't hold.

Comments
11 min read
Routing support tickets with a promise on the errors (and what happens when a new kind of ticket shows up)

Routing support tickets with a promise on the errors (and what happens when a new kind of ticket shows up)

Comments
8 min read
Context engineering vs prompt engineering

Context engineering vs prompt engineering

Comments
3 min read
Model selection is an architecture decision

Model selection is an architecture decision

Comments
2 min read
The Real Post-Mortem: Serving Open LLMs on AWS (SageMaker vLLM vs. Bedrock Custom Models)

The Real Post-Mortem: Serving Open LLMs on AWS (SageMaker vLLM vs. Bedrock Custom Models)

Comments
8 min read
Part 9: Watch and Learn

Part 9: Watch and Learn

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.