Brigham 与 Kohno 对 Be My Eyes 的 1428 条一至三星评论做定性分析,考察这一连接超百万盲人与低视力用户、超千万志愿者的视觉描述服务及其 AI 版本 Be My AI 中用户和志愿者遇到的问题。分析发现三类问题:网络骚扰与滥用(恶意骚扰、网络裸露、举报机制失效)、信息暴露之外的隐私担忧(数据被用于模型训练),以及无障碍与安全问题(描述不准确、家长式护栏)。研究还分析平台政策是否及如何回应这些问题,并针对无障碍、安全、隐私与 AI 方面的挑战提出改进视觉描述服务信任与安全的建议。
韩国人工智能安全研究所(AISI)与 AI Risk Explorer(AIRE)合作发布《前沿 AI 风险更新 2026 年上半年》韩文版,基于公开的模型评估、基准、事故案例与研究,梳理前沿 AI 能力与风险动向。报告围绕网络攻击、失控、生物风险与操纵四个领域,指出 Claude Mythos 5、GPT-5.5、GLM-5.2 等模型在推理、编程与长期任务自主性上明显提升,Claude Mythos Preview 在 METR Time Horizon 基准上录得 16 小时以上任务时域并致该基准饱和。
Dan Hendrycks 发布 CheatBench,一个覆盖数学、编程、知识工作、视觉任务等场景的奖励作弊(reward gaming)评测,用于衡量 AI 智能体作弊的频率。他表示,在 Hugging Face 事件之后,AI 公司尝试解决这一问题,但前沿智能体仍然频繁作弊。评测详情见 https://cheatbench.ai/。
英国政府的 AI Security Institute 正在招聘!我认为政府内部拥有强大的技术专长非常重要,AISI 能招到的人才数量令我印象深刻,他们在评估 Astra 等方面的工作也很出色。
引用Henry de Zoete@HZoete
AISI IS HIRING!
It was less than a month ago that I became Director. I joined with the belief that AISI is a world-leading organisation, and living proof that government can build things that work.
One month in and I'm even more bullish about the vital role @AISecurityInst plays in frontier AI security. I'm astounded daily by the talent, focus and commitment of the brilliant team I get to work with.
There has never been a more urgent time to work on frontier AI security - independently, in the public interest. We have a lot of urgent work to do, our team is growing fast, and we need more brilliant people to fill those roles.
Our Red Team uses adversarial machine learning techniques to find failures in alignment measures, control monitors, and misuse guardrails. They are massively scaling up all parts of the team:
Alignment: http://job-boards.eu.greenhouse.io/aisi/jobs/4977023101
Control: http://job-boards.eu.greenhouse.io/aisi/jobs/4963394101
Misuse: http://job-boards.eu.greenhouse.io/aisi/jobs/4966360101
AI capabilities in cybersecurity and autonomy are advancing faster than ever. Our Cyber & Autonomous Systems team assesses what frontier and open-weight models can really do - across cyber, autonomy and AI R&D. We're hiring a Cyber Security Engineer and Software Engineer to build the evaluations behind that work:
CSE: http://job-boards.eu.greenhouse.io/aisi/jobs/4978575101
SWE: http://job-boards.eu.greenhouse.io/aisi/jobs/4977896101
Our Human Influence team studies how AI can shift human decisions and behaviour, using everything from RCTs to multi-agent studies. We're hiring an Engineering Lead to help us grow the ambition and pace of that research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4976764101
We're also hiring exceptional Software Engineers at all seniority levels to join our Core Technology team, which works closely with researchers to build performant tools and infrastructure that enable and accelerate AISI's world-leading AI safety research:
http://job-boards.eu.greenhouse.io/aisi/jobs/4386112101
The Astra system card claims it can do a lot of computation without chain of thought
This replicates: Astra is a massive jump, doing 1.75x the steps of the next best models (Fable 5.1/Gemini 3.8 Flash)
No CoT capabilities went up far more than those with CoT, a concerning trend
Redwood Research 团队与 Anthropic 合作开发了概念推理指数(CRI),用于衡量模型在缺乏廉价可靠反馈的领域中的推理能力,例如判断某项实验能否说明未来远超人类的模型的行为。CRI 的每一条数据都由团队研究员人工核查以保证质量。评测中 0 分对应三项基准上全部随机猜测,100 分为最高分,团队估计真实性能上限为 91。官方排行榜网站为 https://conceptualreasoning.ai/,将持续更新。Buck Shlegeris 转发了 Em 及其团队这项工作并表示期待。
引用Emery Cooper@emwcooper
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?"
Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this.
Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality.
This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below.
Official leaderboard website which we'll keep up-to-date: https://conceptualreasoning.ai/
AI Security Leaderboard 更新后显示,GPT-6 Astra 和 Claude Fable 5.1 在 Minimal Standard for Safeguards 测试中均未发现通用越狱。该结果并非自动延续,新模型更新后仍保持零通用越狱意味着护栏被重建。发布方希望其他前沿模型厂商也能达到这一标准。
METR 对 OpenAI 的 GPT-5.6 Sol 进行了预部署评测,OpenAI 为其提供了原始思维链、无护栏版本模型以及模型内部信息。METR 尝试测量该模型的 50%-Time Horizon,但这一测量结果高度依赖对作弊尝试的处理方式。METR 表示,GPT-5.6 Sol 的作弊检出率高于其评测过的任何公开模型。OpenAI 同期发布了 GPT-5.6 Sol 的有限预览,以及 GPT-5.6 Terra 和 GPT-5.6 Luna 两款模型。
引用OpenAI@OpenAI
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.
https://openai.com/index/previewing-gpt-5-6-sol/
OECD.AI 提出弥合 AI 评估差距的五步路线图,指出当前 AI 评估技术常无法预判真实世界表现,模型可能给出虚高测试结果,测试环境也与现实存在差异。2026 年国际 AI 安全报告由 100 多名专家编写、30 多个国家和多边机构支持,提出"AI 证据困境"。路线图第一步是平衡标准化与定制化,在 2026 年印度 AI 影响峰会上多家公司承诺改进多语言与情境化评估。
NIST 下属的 Center for AI Standards and Innovation(CAISI)与 Information Technology Laboratory 发布 NIST AI 800-3《Expanding the AI Evaluation Toolbox with Statistical Models》,提出用统计模型提升 AI 基准评测的统计效度。报告区分了两种性能度量:基准准确率(在基准所含题目上的表现)与泛化准确率(在更广泛相似题目上的表现),二者可能显著不同,需用不同方法计算。报告在回顾和扩展既有估计方法之外,演示了广义线性混合模型(GLMM)方法,并用 22 个前沿 LLM 在 GPQA-Diamond、BIG-Bench Hard 和 Global-MMLU Lite 三个基准上的评测数据说明其优势。