DeepSeek-V4 Flash Vision Exp: Official Native Vision vs Real-World Accessibility
- Official release introduces DeepSeek-V4-Flash-Vision-Exp as the first experimental multimodal V4 model with native visual modules, delivering gains...

Created by Cheng Niu
Open‑source and flagship AI model releases, benchmarks, safety notes across LLMs, vision, speech, multimodal
Explore the latest content tracked by AI Model Release Tracker
J-Zero jointly trains a Challenger, Solver, and Judge in a self-play loop, delivering 8.0-point average gains on unverifiable tasks and 4.2 points on...
Two new arXiv works highlight the shift from static vision benchmarks to closed-loop, embodied control.
GPT-5.5 posts the higher aggregate public score at 72.92 versus 51.55, with non-overlapping 90% intervals.
Sora 2 Pro offers up to 1024p resolution and enhanced controls for maximum creative freedom and visual quality on demanding projects. It's built...
Opus 5 posts 43.3% on Frontier-Bench v0.1, more than doubling Opus 4.8’s 18.7% and topping Fable 5’s 33.7% at identical $5/$25 pricing.
Newly released models dominating the Terminal-Bench 2.1 leaderboard may reflect benchmark-specific optimization rather than genuine frontier terminal reasoning. Top 20 models were retested on TB-fn, which reformulates the same 89 tasks.
Video generators create visually convincing clips yet fail to reproduce the true distribution of physical outcomes when rolled out repeatedly.
-...
GPT-5.6 Sol leads Big Finance Bench at 0.530 on complex financial reasoning tasks, ahead of two other OpenAI entries. All share 1.1M context windows with API costs ranging $0.20–$5.00.
公开资料中两款模型无共同基准结果,无法得出质量优劣结论。只能参考GPT-4.1已披露的单项指标(如Coding 54.6、Knowledge 66.3)和成本数据(1K输入+500输出约$0.006),Lyra Mini多类数据未公布或缺失API费率。
MiniMax H3 generates AI videos from text, images, and reference clips at up to 2K with synchronized sound. The browser demo at https://minimax-h3.im/ enables fast experiments with short-form concepts.
Meta's SIGReg method for LeVJEPA delivers a 20x FLOP efficiency improvement over VJEPA1/2 pretraining by using only a simple sigreg + prediction loss,...
Tencent released the open-source Hy3 model with 770B parameters on Aug 28, 2026, targeting coding, research, and financial analysis. No benchmarks or deployment details were provided in the announcement.
Could a specialized LLM meaningfully cut documentation workload for knee osteoarthritis and osteoporosis in Indonesian primary care?
Gemini 3.5 Flash posts a higher public score estimate (64.73 vs 51.07) yet the 90% intervals overlap, so the difference is only a lead, not a settled...
Interfaze Beta leads the VoxPopuli WER benchmark at 2.4%, the sole confirmed result among evaluated models. This standardized word-error metric isolates speech-recognition gains without relying on general rankings.