Kimi K3 视觉分数值得每样本成本吗
Kimi K3 六项视觉基准平均 66.5%,OCR 排名第5,每样本成本约 0.011 美元。
从榜单到采购:这一性能与价格组合是否足够吸引真实应用?

Created by Yizengxiong Zhu
Cutting‑edge NLP and computer vision research, industry updates, and AI safety policy coverage
Explore the latest content tracked by Vision & Language Pulse
Kimi K3 六项视觉基准平均 66.5%,OCR 排名第5,每样本成本约 0.011 美元。
从榜单到采购:这一性能与价格组合是否足够吸引真实应用?
视觉语言模型需同时掌握排序与生成两项核心能力,仅靠单一目标训练无法兼得。 基础训练配方正是结合 InfoNCE 对比损失(拉近匹配图文对、区分负样本)与图像条件自回归语言建模损失(训练解码器生成描述),二者联合优化即可在单次预训练中同时习得两种技能。
本周AI Agent论文聚焦Skill Lift、JIT-Agent及代码化上下文管理。
When robot trajectories resist scaling due to cost and sparsity, representation-centric continued pre-training proves more effective than simply...
Multimodal agents now face a sharper test: dexterous visual tool use—inferring precise parameters from visual evidence in closed-loop execution. On...
工具调用型Agent的安全防护正从事后拦截转向系统级架构。StepGuard实现步骤级护栏,可在工具执行前审计动作,AgentDojo攻击成功率下降77.3%,效用仅降2.8%。
LMSM借鉴Linux安全模块,将安全后端、版本化策略与门控分离,实现多规则组合与运行时灵活控制。
两者共同指向:逐步行为护栏与类OS策略控制的结合,将成为Agent安全的主流方向。
8月29日,欧盟AI Office向OpenAI、Anthropic、Google等通用AI模型提供商发出首批正式信息请求,聚焦模型安全、外部评估及市场监测。
此举距8月2日义务生效仅四周,源于夏季多起模型失控事件。回复不实或缺席可面临最高1500万欧元或全球营业额3%罚款。
训练数据摘要要求同时针对未公开版权信息的提供商,标志透明度与风险管理进入强制记录阶段。
A single-lineage prompt optimizer matched GEPA on instruction-following benchmarks while using similar or lower rollout budgets, with no candidate...
ContextPilot equips long-horizon agents with an expanded toolset—planning, long-term memory, and soft offloading—plus fine-grained RL that scores...
长链路推理代理因全注意力机制需完整保留轨迹,导致 hardest problems 内存压力巨大。 斯坦福新工作聚焦高效测试时缩放,在不保留完整轨迹的前提下维持 agent 性能。
视频生成模型可 repurposed 为统一几何估计框架,将深度与表面法向预测视为 next-frames 任务,GeoNeXt 继承视频先验实现 image-geometry 联合建模,仅用少量数据即超越 prior generative 方法,并在多基准上 rival 需 100x 数据训练的 discriminative SOTA。
Vision-language models gain robust representations via novel techniques:
OpenAI's GPT-5.4 mini and nano deliver serious power in smaller footprints, validated by key benchmarks for edge and developer use:
NVIDIA's GTC 2026 open model push spans agentic-physical-healthcare AI:
The Emerging Science of Machine Learning Benchmarks book is buzzing with 35 points on Hacker News, spotlighting the evolving science behind ML evaluation standards.
Trend spotlight: Robustness challenges intensify in vision-language AI.