DEV Community

Deep Learning

This tag is for discussing, sharing articles, and asking questions primarily on deep learning - a subfield of machine learning.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
UniEvo-VL: On-Policy Self-Distillation for Multimodal Image Generation

UniEvo-VL: On-Policy Self-Distillation for Multimodal Image Generation

Comments
5 min read
RoPE vs sinusoidal positional encoding, measured: 55 logits of drift against 0.0005

RoPE vs sinusoidal positional encoding, measured: 55 logits of drift against 0.0005

Comments
7 min read
Why AI Lies Without Knowing It’s Lying | Victor Amit

Why AI Lies Without Knowing It’s Lying | Victor Amit

1
Comments
21 min read
REAL-Q: How Dynamic Gradient Descent Fixes the Core Flaw in LLM Quantization

REAL-Q: How Dynamic Gradient Descent Fixes the Core Flaw in LLM Quantization

Comments
5 min read
I Built a Tiny Neural Network From Scratch in Python — No PyTorch

I Built a Tiny Neural Network From Scratch in Python — No PyTorch

1
Comments
6 min read
We Tried ISO-AdamW. AdamW Kept Its Job.

We Tried ISO-AdamW. AdamW Kept Its Job.

1
Comments
11 min read
Bonsai 2 27B: How Ternary Weights and a Hadamard Rotation Compress a 27B Model to 5.9 GB

Bonsai 2 27B: How Ternary Weights and a Hadamard Rotation Compress a 27B Model to 5.9 GB

Comments 1
5 min read
What I Learned Building a Two-Image AI Virtual Try-On Workflow for Holiday Outfits

What I Learned Building a Two-Image AI Virtual Try-On Workflow for Holiday Outfits

Comments 1
6 min read
nn.Module Explained: The Same Model Built with Raw Tensors and with nn.Module

nn.Module Explained: The Same Model Built with Raw Tensors and with nn.Module

Comments
8 min read
Ornith-1.0-9B vs. Qwen3.5 vs. Gemma4: A Local LLM Battle Royale

Ornith-1.0-9B vs. Qwen3.5 vs. Gemma4: A Local LLM Battle Royale

1
Comments
5 min read
RL 4: Early heuristics and the birth of Temporal Difference learning (1959–1968)

RL 4: Early heuristics and the birth of Temporal Difference learning (1959–1968)

Comments
15 min read
An Overview of Basic Gradient Descent Optimization Algorithms

An Overview of Basic Gradient Descent Optimization Algorithms

Comments
6 min read
Kimi Linear: How Moonshot AI Built a Hybrid Attention Architecture That Beats Full Attention

Kimi Linear: How Moonshot AI Built a Hybrid Attention Architecture That Beats Full Attention

Comments
5 min read
How Multi-Agent AI Discovered a New Enzyme System in Phage DNA

How Multi-Agent AI Discovered a New Enzyme System in Phage DNA

Comments
4 min read
Recognizing ASL Letters with CNN and Ensemble Learning

Recognizing ASL Letters with CNN and Ensemble Learning

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.