Engineering
Official technical announcement and publication from Hugging Face covering TRL v1.0: Post-Training Library Built to Move with the Field.
Official LMSYS Chatbot Arena release and benchmark update covering Highlights of SGLang at NVIDIA GTC 2026.
Artificial intelligence (AI) predictions are increasingly used to inform human decisions. Here, using a behavioral implementation of the classic Newcomb's paradox in 1,305 participants, we show that AI predictions can also shape the reasoning people use to make a decision. In this paradigm, perceived predictive authority can alter how people reason about their future actions, leading them to forgo
Despite there now being more than 1,000 FDA-authorised AI medical devices, formal equity assessments -- whether model performance is uniform across patient subgroups -- are rare. Here, we evaluate the equity of 18 open-source brain tumour segmentation models across 648 glioma patients from two independent datasets (n = 11,664 model inferences) along distinct univariate, Bayesian multivariate, spat
AI for Disaster Response in Asia: OpenAI Workshop with Gates Foundation
Google DeepMind is transforming the mouse pointer into a context-aware AI partner. Move beyond the friction of traditional prompting with intuitive AI collaboration in Chrome and beyond.
AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accepted benchmark exists for evaluating their performance. We introduce PeopleSearchBench, an open-source benchmark that compares four people search platforms on 119 real-world queries across four use cases: corporate recruiting, B2B sales prospecting, expert searc
Learn how STADLER uses ChatGPT to transform knowledge work, saving time and accelerating productivity across 650 employees.
Official technical announcement and publication from Hugging Face covering Liberate your OpenClaw.
March 27, 2026
Official a16z (Andreessen Horowitz) technical update and publication covering AI, Supply Chains, and the Future of Economic Power.
Official a16z (Andreessen Horowitz) technical update and publication covering What’s Missing Between LLMs and AGI – Vishal Misra & Martin Casado.
Official a16z (Andreessen Horowitz) technical update and publication covering AI Startups vs. Big Chatbots — With Olivia Moore.
Official a16z (Andreessen Horowitz) technical update and publication covering Emil Michael: Iran, Anthropic and the Future of AI at the Pentagon.
Official a16z (Andreessen Horowitz) technical update and publication covering What It Takes to Clear a Million Crimes a Year with Flock Safety’s CEO.
Official a16z (Andreessen Horowitz) technical update and publication covering Atlassian CEO on the SaaS Apocalypse, AI Agents & What Comes Next.
Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.
Introducing Cohere Transcribe: a new state-of-the-art in open-source speech recognition
As context windows grow, LLM performance degrades in unexpected ways. We show how a "Divide & Conquer" framework — breaking long documents into parallel chunks with a planner, workers, and manager — lets smaller models like Llama-3-70B and Qwen-72B outperform GPT-4o single-shot.
March 26, 2026
Google DeepMind researches AI's harmful manipulation risks across areas like finance and health, leading to new safety measures.
Introducing Lyria 3 Pro, which unlocks longer tracks with structural awareness. We’re also bringing Lyria to more Google products and surfaces.
Learn how OpenAI’s Model Spec serves as a public framework for model behavior, balancing safety, user freedom, and accountability as AI systems advance.
Human-Computer Interaction and Visualization
OpenAI launches a Safety Bug Bounty program to identify AI abuse and safety risks, including agentic vulnerabilities, prompt injection, and data exfiltration.
Research & Technology
Sierra is reimagining software for the agent era—where you simply describe the outcome, and intelligent agents build, execute, and continuously improve the work for you. Meet Ghostwriter, the agent that creates and optimizes other agents, turning your ideas into production-ready customer experiences without clicks, code, or complexity.
Research & Technology
Research & Technology
Official LMSYS Chatbot Arena release and benchmark update covering Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE Deployments.
Algorithms & Theory
Algorithms & Theory
OpenAI releases prompt-based teen safety policies for developers using gpt-oss-safeguard, helping moderate age-specific risks in AI systems.
The OpenAI Foundation announces plans to invest at least $1 billion in curing diseases, economic opportunity, AI resilience, and community programs.
ChatGPT introduces richer, visually immersive shopping powered by the Agentic Commerce Protocol, enabling product discovery, side-by-side comparisons, and merchant integration.
Official technical announcement and publication from Hugging Face covering A New Framework for Evaluating Voice Agents (EVA).
Official Cohere release and benchmark update covering Hardware-aware dynamic speculative decoding.
In natural conversation, the gap between one person finishing a sentence and the other starting to respond averages around 200 milliseconds. For voice agents this is the target to match.
Posted on March 24, 2026
Voxtral TTS
Large Language Models (LLMs) are deployed in high-stakes settings but can show demographic, gender, and geographic biases that undermine fairness and trust. Prior debiasing methods, including embedding-space projections, prompt-based steering, and causal interventions, often act at a single stage of the pipeline, resulting in incomplete mitigation and brittle utility trade-offs under distribution
To address the novel safety challenges posed by a state-of-the-art video model as well as a new social creation platform, we’ve built Sora 2 and the Sora app with safety at the foundation. Our approach is anchored in concrete protections.
Official Index Ventures technical update and publication covering Read more Opens in a new window..
Official Decagon technical update and publication covering What we’ve learned about designing AI-ready CX teams.
Official technical announcement and publication from Hugging Face covering Build a Domain-Specific Embedding Model in Under a Day.
Devin can now schedule its own recurring sessions. Run a task once, and if it goes well, tell Devin to keep doing it. It maintains state between runs, so each session picks up where the last one left off.
March 20, 2026
As large language models (LLMs) are increasingly deployed as automated graders in educational settings, concerns about fairness and bias in their evaluations have become critical. This study investigates whether LLMs exhibit implicit grading bias based on writing style when the underlying content correctness remains constant. We constructed a controlled dataset of 180 student responses across thre
How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI safety safeguards.
Accelerates Codex growth to power the next generation of Python developer tools
Devin can now break down large tasks and delegate them to a team of managed Devins, with each running in its own isolated VM in parallel.
March 19, 2026
Posted on March 19, 2026
Together AI expands fine-tuning with native support for tool call, reasoning, and vision-language models, plus 100B+ model training, up to 6× higher throughput, and job cost and ETA estimates.
Health & Bioscience
Incident management is essential to maintain the reliability and availability of cloud computing services. Cloud vendors typically disclose incident reports to the public, summarizing the failures and recovery process to help minimize their impact. However, such reports are often lengthy and unstructured, making them difficult to understand, analyze, and use for long-term dependability improvement
Health & Bioscience
Official technical announcement and publication from Hugging Face covering State of Open Source on Hugging Face: Spring 2026.
We’re introducing a framework to measure progress toward AGI, and launching a Kaggle hackathon to build the relevant evaluations.
Official technical announcement and publication from Hugging Face covering Holotron-12B - High Throughput Computer Use Agent.
OpenAI Japan announces the Japan Teen Safety Blueprint, introducing stronger age protections, parental controls, and well-being safeguards for teens using generative AI.
GPT-5.4 mini and nano are smaller, faster versions of GPT-5.4 optimized for coding, tool use, multimodal reasoning, and high-volume API and sub-agent workloads.
Mar 17, 2026
Meet Mamba-3: the SSM built for inference. Faster than Transformers at decode, stronger than Mamba-2, and open-source from day one.
New research shows Americans send nearly 3 million daily messages to ChatGPT asking about compensation and earnings, helping close the wage information gap.
Official Mistral AI technical update and publication covering Introducing Forge.
Official LMSYS Chatbot Arena release and benchmark update covering ROCm Support for Miles: Large-Scale RL Post-Training on AMD Instinct™ GPUs.
Mistral Small 4
Education Innovation
Together AI arrives at NVIDIA GTC 2026 with new launches in inference, agents, voice AI, and open models — plus technical sessions from its research and engineering leaders.
A deep dive into why Codex Security doesn’t rely on traditional SAST, instead using AI-driven constraint reasoning and validation to find real vulnerabilities with fewer false positives.
Cursor built a fleet of security agents to solve a familiar frustration
Official Mistral AI technical update and publication covering Leanstral: Open-Source foundation for trustworthy vibe-coding.
Official Mistral AI technical update and publication covering Mistral AI partners with NVIDIA to accelerate open frontier models.
<!-- twitter --> <meta name="twitter:title" content="Identifying Interactions at Scale for LLMs" /> <meta name="twitter:card" content="summary_large_image" /> <meta name="twitter:image" content="https://bair.berkeley.edu/static/blog/spex/teaser.png" /> <meta name="keywords" content="" /> <meta name="description" content="The BAIR Blog" /> <meta name="author" content="Landon Butler, Justin Singh Ka
March 13, 2026
Climate & Sustainability
The department is saddened to announce the death of Emeritus Professor Sir Tony Hoare.
Climate & Sustainability
Build real-time voice agents on Together AI with co-located STT, LLM, and TTS infrastructure, native Deepgram and Cartesia support, and end-to-end latency under 500ms.
Mar 12, 2026
Generative AI
How ChatGPT defends against prompt injection and social engineering by constraining risky actions and protecting sensitive data in agent workflows.
How OpenAI built an agent runtime using the Responses API, shell tool, and hosted containers to run secure, scalable agents with files, tools, and state.
Official technical announcement and publication from OpenAI covering Rakuten fixes issues twice as fast with Codex.
Wayfair uses OpenAI models to improve ecommerce support and product catalog accuracy, automating ticket triage and enhancing millions of product attributes at scale.
NVIDIA Nemotron 3 Super is now available on Together AI Dedicated Inference, delivering efficient multi-agent reasoning, a 1M-token context window, and production-grade deployment on managed infrastructure.
Official LMSYS Chatbot Arena release and benchmark update covering SGLang Adds Day-0 Support for NVIDIA Nemotron 3 Super for building High-Efficiency Multi-Agent Systems.
Official Mistral AI technical update and publication covering Rails testing on autopilot: Building an agent that writes what developers won't.
IH-Challenge trains models to prioritize trusted instructions, improving instruction hierarchy, safety steerability, and resistance to prompt injection attacks.
ChatGPT introduces interactive visual explanations for math and science, helping students explore formulas, variables, and concepts in real time.
Together GPU Clusters now include built-in autoscaling, RBAC, full-stack observability, and self-healing node repair—giving teams production-ready GPU infrastructure that scales efficiently, stays resilient, and supports shared enterprise workloads.
Official technical announcement and publication from Hugging Face covering Introducing Storage Buckets on the Hugging Face Hub.
Official technical announcement and publication from Hugging Face covering Keep the Tokens Flowing: Lessons from 16 Open-Source RL Libraries.
Mar 10, 2026
Ten years since AlphaGo, we explore how it is catalyzing scientific discovery and paving a path to AGI.
The digital markets act (DMA) regulates very large digital platforms like Meta's Facebook or Apple's iOS with the goal to promote fairness, contestability (of market power) and user choice. From a system design or broader technical perspective, the implications of the DMA have not been studied so far. Using systematic methods from qualitative coding and thematic analysis, we investigate the DMA fr
OpenAI is acquiring Promptfoo, an AI security platform that helps enterprises identify and remediate vulnerabilities in AI systems during development.
Official technical announcement and publication from Hugging Face covering Ulysses Sequence Parallelism: Training with Million-Token Contexts.
6584 articles sourced historically · 100 per page