OpenAI Moves Under Dots and Question Marks
Stratechery's "Dots and Question Marks" analysis examines OpenAI's actions, framing key developments that could shape persistent AI agents and the broader competitive landscape.

Created by weiqun zou
Latest news, benchmarks, and use‑case guides for AI coding assistants and autonomous agents
Explore the latest content tracked by AI Coding Tools Digest
Stratechery's "Dots and Question Marks" analysis examines OpenAI's actions, framing key developments that could shape persistent AI agents and the broader competitive landscape.
Unprompted bug hunting remains far harder for agents than standard issue-resolution benchmarks because models receive a repo at a past commit with...
MIT and Sakana AI's SIFT framework replaces expensive full benchmarks with cheap LLM pairwise judgments, enabling parallel tree search across agent...
Ambiguous terms like "available" let agents generate fluent but conflicting interpretations across sales, manufacturing, and suppliers, creating...
openai-codex models while injecting service_tier: "priority" at the boundaryDevelopers praise GLM-5.3 Flash for minor fixes and post-planning implementation, calling it a reliable workhorse after switching from heavier...
AI coding tools produce working code fast, but correctness requires deliberate review.
Apply this repeatable workflow:
Microsoft's job post reveals aggressive internal investment in agentic developer tools through targeted hiring.
A two-tier architecture lets AI agents sustain multi-day workflows without losing critical context.
Agent security must rely on external authorization, isolation, and policy controls rather than model behavior alone. NVIDIA's Open Agent Safety...
A new Pi extension delivers zero-config Codex quota visibility directly in the statusline.
GRPO's uniform token-level advantages fail to distinguish decisive steps from the rest of a trajectory. ProVer instead uses an LLM judge to identify...
AI tools let developers complete 21% more tasks, but review time jumps 91% as they juggle 47% more simultaneous workstreams. The real bottleneck has moved from writing code to trusting it.
Retrieval benchmarks for agents gain real value when built on messy company knowledge instead of clean synthetics. Discussions stress publishing extraction pipelines to enable reproducibility across organizations.
AI slots into GitLab pipelines as ordinary jobs without altering the core architecture of stages, jobs, and runners.
GraphForge tackles the core gap in training repository-scale agents: the lack of tasks built on real files with coordinated tools and verifiable...
Apple is adding stricter controls to macOS Full Disk Access after AI agents like Meta's Muse exposed users' messages, files, and browsing history...
This hands-on lab moves past chatbot tutorials to teach production-grade agent systems.