Thinking Machines Lab unveiled interaction models with 0.4s latency and simultaneous audio-video-text. Here’s how the dual-model architecture works, benchmarks vs GPT and Gemini, and what it means for enterprise.
Browsing: full-duplex AI
An interaction model is an AI system that listens, watches, and speaks in continuous 200-millisecond beats – instead of waiting for your turn to end before thinking. A plain-English guide to what Thinking Machines just announced, with benchmarks, architecture, demos, and what it means for your AI stack.
