Episode Details
Back to Episodes
AI Papers - 2026-06-26
Published 1 month, 1 week ago
Description
Today's papers:
- The Capability Frontier: Benchmarks Miss 82% of Model Performance: https://arxiv.org/abs/2606.26836v1
- Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection: https://arxiv.org/abs/2606.26552v1
- auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation: https://arxiv.org/abs/2606.26460v1
- TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs: https://arxiv.org/abs/2606.26029v1
- TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference: https://arxiv.org/abs/2606.27161v1
This podcast is from Colin Davis (colin-davis.com) using Claude & Elevenlabs.