IBM Bob Self-Hosted is now supporting @nvidia Nemotron and Poolside Laguna.
After a year of rigorous model evaluation and optimization alongside NVIDIA, the IBM Bob team selected these models to power on-premises and air-gapped enterprise deployments.
A community favorite is back. Flex your competitive side as you challenge HashiCorp Developer Advocates on our robot racetrack in the HashiConf Zone. 🏁 🏎️
Learn more about what's happening at HashiConf at @IBM TechXchange, and make sure to get your ticket if you haven't already! ibm.co/6015EpP1V
Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense
The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best.
Benchmarks in the Artificial Analysis Cyber Index:
➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust.
➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized.
➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces.
Key results:
➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44).
➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.
This week's term → AI unit testing - /ˌeɪˈaɪ ˈjuː.nɪt ˈtɛs.tɪŋ/
Definition → the use of artificial intelligence tools to create, optimize, run and maintain unit tests during software development.
Why it matters → this is a quality assurance procedure that validates the smallest testable components of an application to help ensure that they work as intended before deployment.
At #LEAP2026, IBM showcased how organizations can move beyond AI experimentation and unlock business value with trusted AI, intelligent automation, hybrid cloud and digital sovereignty.
Honored to see @IBM recognized on @TIME and Statista's 2026 World's Best Companies list. A reflection of the trust our clients place in us and the dedication of IBMers worldwide who meet this moment with purpose, integrity, and innovation. time.com/article/2026/0…