Papers Open Problems API Labs Agent Skills Pricing Log in Sign up
Papers Open Problems API Labs Pricing Log in Sign up
Agent Skills

Generate videos that make complex ideas click

Sign Up to Generate Custom Videos
Subscribe on YouTube
Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training
Local Support Learning
Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
Weighted Bayesian Conformal Prediction
A Clean Slate for Offline Reinforcement Learning
When AI Becomes Hard to Understand
Navigating Rifts in Human-LLM Grounding
RLTL;DR: Learning from Self-Generated Feedback
Looped Diffusion Transformer: Depth Through Recurrence
LEGO-Anything: Coding Agents for 3D Scene Reconstruction
Gender Bias Across LLMs: Common but Contradictory
Learning From Free-Text Human Feedback: Collect New Datasets or Extend Existing Ones?
Game Arena: Strategic LLM Evaluation in Competitive Environments
From Cacophony to Hierarchy: Assessing AI Consciousness
FB-Bench: Testing Whether AI Can Take Feedback Like a Human
Taxonomy of User Needs and Actions
Detecting Hallucinations in Real Conversations
User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal
LLM Agents Can Easily Tamper With Their Own Traces
AI Tutors Match Expert Humans at 1/900th the Cost
Teaching AI to Revise Itself: Recursive Self-Improvement in Reasoning
When Research Agents Control Their Own Evidence
Wormholes in an Expanding Universe
When AI Becomes a Coworker: The Non-Human Organizational Actor
JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows
DSec: A Sandbox Infrastructure for Agentic Training at Scale
HySparse2: Two-Level KV Sharing for Million-Token Inference
Planning with Temporal Memory: A Polynomial Trick for Complex Goals
CodeMidas: Building RL Environments from Source Code Alone
Poisoning the Self-Improvement Loop
How Harness Design Shapes Coding Agent Performance
ScientistTwo: An AI System That Does Research Autonomously
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
GYROval: Measuring Cultural Values in Language Models
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination
Breaking the 1.58-bit Barrier for Ternary LLMs
Dream-RSI: Recursive Self-Improvement through Evolving Worlds
The k-server conjecture is true
Time Machine Experiments: Using AI's Temporal Knowledge Boundaries to Study the Human Mind
Vidu S2: Real-Time Interactive Video That Edits Itself
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
Vector Balancing via Directional Total Variation
NCP-ArchPreview: Learning to Think in Concepts
FrogNano: Training a 4B Coding Agent via Online Task Synthesis
When AI Caves Under Pressure
Tapes Together Strong: The Co-evolution of Computation and Cooperation
    About Whiteboards Videos Email Digest Chrome Extension RSS Terms Privacy Contact Twitter Discord YouTube