At Sierra Summit, we unveiled the future of AI agents: a single agent that works across every channel, learns from every interaction, and builds real customer relationships. With eight new products, including Agent Studio 2.0 and the Agent Data Platform, Sierra is turning the promise of AI-powered customer experience into reality.
Sierra’s first-of-its-kind Agent Data Platform (ADP) gives agents memory, context, and intelligence — turning every interaction into a human, personalized experience. Powered by Agent OS, ADP helps companies move from answering questions to anticipating needs.
Posted on November 5, 2025
General Science
Understanding how to evaluate and benchmark Large Language Models (LLMS). Test, compare, and understand LLMs.
Together AI launches the fastest voice AI stack: streaming Whisper STT, serverless open-source TTS (Orpheus & Kokoro), and Voxtral transcription. Sub-second latency for production voice agents.
Codemaps is meant to offer a shared understanding of a system between humans and AI, enabling your AI to teach you about the code you are looking at quickly and elegantly. A codemap can be generated about any system or snippet to illuminate its code paths, helping users learn and recall. Codemaps allows AI to be a partner that explains code in an accurate and consistent way, rather than generating
OpenAI introduces IndQA, a new benchmark for evaluating AI systems in Indian languages. Built with domain experts, IndQA tests cultural understanding and reasoning across 12 languages and 10 knowledge areas.
OpenAI and AWS have entered a multi-year, $38 billion partnership to scale advanced AI workloads. AWS will provide world-class infrastructure and compute capacity to power OpenAI’s next generation of models.
November 03, 2025
November 03, 2025
Official LMSYS Chatbot Arena release and benchmark update covering Optimizing GPT-OSS on NVIDIA DGX Spark: Getting the Most Out of Your Spark.
<!-- twitter --> <meta name="twitter:title" content="RL without TD learning" /> <meta name="twitter:card" content="summary_large_image" /> <meta name="twitter:image" content="https://bair.berkeley.edu/static/blog/rl-without-td-learning/teaser.png" /> <meta name="keywords" content="" /> <meta name="description" content="The BAIR Blog" /> <meta name="author" content="Seohong Park" /> <p>In this post
As large language models become increasingly capable of generating code, evaluating their performance remains a complex and evolving challenge. Existing benchmarks primarily focus on functional correctness, overlooking the diversity of real-world coding tasks and developer expectations. To this end, we introduce a multi-language benchmark that evaluates LLM instruction-following capabilities and i
Climate & Sustainability
OpenAI is expanding Stargate to Michigan with a new one-gigawatt campus that strengthens America’s AI infrastructure. The project will create jobs, drive investment, and support economic growth across the Midwest.
OpenAI introduces Aardvark, an AI-powered security researcher that autonomously finds, validates, and helps fix software vulnerabilities at scale. The system is in private beta—sign up to join early testing.
Generative AI
Official technical announcement and publication from Hugging Face covering Aligning to What? Rethinking Agent Generalization in MiniMax M2.
A deep dive into OWL, the new architecture powering ChatGPT Atlas—decoupling Chromium, enabling fast startup, rich UI, and agentic browsing with ChatGPT.
Official Harvey AI technical update and publication covering Al Hounsell on Driving Legal Transformation That Lasts.
Official Decagon technical update and publication covering From support to growth: The next evolution of AI agents.
Generative AI
The initiative brings together some of the world's most prestigious research institutions to pioneer the use of AI in mathematical research.
Official technical announcement and publication from Hugging Face covering On the Shifting Global Compute Landscape.
Today we’re releasing SWE-1.5, the latest in our family of models optimized for software engineering. It is a frontier-size model with hundreds of billions of parameters that achieves near-SOTA coding performance. It also sets a new standard for speed: we partnered with Cerebras to serve it at up to 950 tok/s – 6x faster than Haiku 4.5 and 13x faster than Sonnet 4.5. SWE-1.5 is now available in Wi
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are two open-weight reasoning models post-trained from the gpt-oss models and trained to reason from a provided policy in order to label content under that policy. In this report, we describe gpt-oss-safeguard’s capabilities and provide our baseline safety evaluations on the gpt-oss-safeguard models, using the underlying gpt-oss models as a baseline
Official technical announcement and publication from Hugging Face covering Building a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac.
OpenAI introduces gpt-oss-safeguard—open-weight reasoning models for safety classification that let developers apply and iterate on custom policies.
Official LMSYS Chatbot Arena release and benchmark update covering SGLang-Jax: An Open-Source Solution for Native TPU Inference.
Official technical announcement and publication from Hugging Face covering How to Build a Healthcare Robot from Simulation to Deployment with NVIDIA Isaac for Healthcare.
DNP rolled out ChatGPT Enterprise across ten core departments, achieving 95% faster patent research, 10x processing volume, 87% automation, and 70% knowledge reuse in three months.
Official technical announcement and publication from Hugging Face covering Granite 4.0 Nano: Just how small can you go?.
Doppel uses GPT-5 and reinforcement fine-tuning to stop deepfake and impersonation attacks, cutting analyst workloads by 80% and reducing response times from hours to minutes.
Microsoft and OpenAI sign a new agreement that strengthens its long-term partnership, expands innovation, and ensures responsible AI progress.
OpenAI’s recapitalization strengthens mission-focused governance, expanding resources to ensure AI benefits everyone while advancing innovation responsibly.
Engineering teams dread migrating from .NET Framework to .NET Core. However, autonomous coding agents are rapidly changing the previously known timelines. What once took months, teams are now finishing in as little as two weeks — with Devin.
Official Groq (LPUs) technical update and publication covering Groq Powers HUMAIN One, a Real-Time AI Operating System for Enterprise.
Official technical announcement and publication from Hugging Face covering Voice Cloning with Consent.
Test AI agents in the real world with Collinear TraitMix and Together Evals: dynamic persona simulations, multi-turn dialogs, and LLM-as-judge scoring.
Generative AI
Meeting the demands of the Intelligence Age will require strategic investment in energy and infrastructure. OpenAI’s submission to the White House details how expanding capacity and workforce readiness can sustain U.S. leadership in AI and economic growth.
OpenAI collaborated with 170+ mental health experts to improve ChatGPT’s ability to recognize distress, respond empathetically, and guide users toward real-world support—reducing unsafe responses by up to 80%. Learn how we’re making ChatGPT safer and more supportive in sensitive moments.
This system card details GPT-5’s improvements in handling sensitive conversations, including new benchmarks for emotional reliance, mental health, and jailbreak resistance.
We’re discontinuing free API credits due to widespread abuse and spam signups.
Steuerrecht.com uses ChatGPT Business to streamline legal workflows, automate tax research, and deliver faster, client-ready analysis for law firms.
Official technical announcement and publication from Hugging Face covering huggingface_hub v1.0: Five Years of Building the Foundation of Open Machine Learning.
We’re discontinuing free API credits due to widespread abuse and spam signups.
Official technical announcement and publication from Hugging Face covering Streaming datasets: 100x More Efficient.
October 27, 2025
Introducing T5Gemma, a new collection of encoder-decoder LLMs.
We’re announcing new multimodal models in the MedGemma collection, our most capable open models for health AI development.
Gemma 3n is designed for the developer community that helped shape Gemma.
Gemini 2.5 Flash-Lite, previously in preview, is now stable and generally available. This cost-efficient model provides high quality in a small size, and includes 2.5 family features like a 1 million-token context window and multimodality.
We partnered with Darren Aronofsky, Eliza McNitt and a team of more than 200 people to make a film using Veo and live-action filmmaking.
New AI model integrates petabytes of Earth observation data to generate a unified data representation that revolutionizes global mapping and monitoring
He joins Professor Marina Jirotka, Institute Director, in leading the internationally recognised centre of excellence focused on responsible technology.
New experimental AI tool helps people explore the context and origin of images seen online.
The International Mathematical Olympiad (“IMO”) is the world’s most prestigious competition for young mathematicians, and has been held annually since 1959. Each country taking part is represented by six elite, pre-university mathematicians who compete to solve six exceptionally difficult problems in algebra, combinatorics, geometry, and number theory.
Introducing the first model for contextualizing ancient inscriptions, designed to help historians better interpret, attribute and restore fragmentary texts.
Genie 3 can generate dynamic worlds that you can navigate in real time at 24 frames per second, retaining consistency for a few minutes at a resolution of 720p.
Our new Perch model helps conservationists analyze audio faster to protect endangered species, from Hawaiian honeycreepers to coral reefs.
Using AI to perceive the universe in greater depth
Gemini 2.5 Deep Think achieves breakthrough performance at the world’s most prestigious computer programming competition, demonstrating a profound leap in abstract problem solving.
Our new method could help mathematicians leverage AI techniques to tackle long-standing challenges in mathematics, physics and engineering.
Official technical announcement and publication from Hugging Face covering LeRobot v0.4.0: Supercharging OSS Robot Learning.
Official Mistral AI technical update and publication covering Introducing Mistral AI Studio..
The contest is part of the International Collegiate Programming Contest series.
OpenAI has acquired Software Applications Incorporated, maker of Sky—a natural language interface for Mac that brings AI directly into your desktop experience. Together, we’re integrating Sky’s deep macOS capabilities into ChatGPT to make AI more intuitive, contextual, and action-oriented.
Consensus uses GPT-5 and OpenAI’s Responses API to power a multi-agent research assistant that reads, analyzes, and synthesizes evidence in minutes—helping over 8 million researchers accelerate scientific discovery.
Climate & Sustainability
OpenAI's Korea Economic Blueprint outlines how South Korea can scale trusted AI through sovereign capabilities and strategic partnerships to drive growth.
Company knowledge brings context from your apps into ChatGPT for answers specific to your business, with clear citations, security, privacy, and admin controls. Available now for Business, Enterprise, and Edu users.
Official technical announcement and publication from Hugging Face covering Building the Open Agent Ecosystem Together: Introducing OpenEnv.
OpenAI expands its UK partnership with a new Ministry of Justice agreement, bringing ChatGPT to civil servants. It also introduces UK data residency for ChatGPT Enterprise, ChatGPT Edu, and the API Platform to support trusted and secure AI adoption.
Official technical announcement and publication from Google Research covering A verifiable quantum advantage.
Oct 22, 2025
Official technical announcement and publication from Hugging Face covering Hugging Face and VirusTotal collaborate to strengthen AI security.
OpenAI’s Japan Economic Blueprint outlines how Japan can harness AI to boost innovation, strengthen competitiveness, and enable sustainable, inclusive growth.
ReasonIF finds frontier LRMs fail to follow reasoning instructions >75% of the time; introduces a benchmark across languages, formatting, and length.
Official technical announcement and publication from Hugging Face covering Sentence Transformers is joining Hugging Face!.
Official LMSYS Chatbot Arena release and benchmark update covering Accelerating Hybrid Inference in SGLang with KTransformers CPU Kernels.
ChatGPT will no longer be available on WhatsApp after January 15, 2026. Learn how to link your ChatGPT account and continue your conversations across devices.
Together AI adds 40+ image & video models, including Sora 2 and Veo 3, to build end-to-end multimodal apps with unified OpenAI-compatible APIs and transparent pricing.
This post introduced ChatGPT Atlas, a browser with ChatGPT built in. Atlas has since been deprecated.
Official technical announcement and publication from Hugging Face covering Unlock the power of images with AI Sheets.
Soniox v3 is our most advanced speech AI yet — designed for accuracy, fluency, and speed at scale.
Soniox v3 is our most advanced speech AI yet — designed for accuracy, fluency, and speed at scale.
Official technical announcement and publication from Hugging Face covering Supercharge your OCR Pipelines with Open Models.
Generative AI
DPhil student Tiffany Horter discusses why we need humans making mistakes when designing AI systems.
General Science
Algorithms & Theory
At Meta, we are constantly pushing the boundaries of LLM inference systems to power applications such as the Meta AI App. We’re sharing how we developed and implemented advanced parallelism techniques to optimize key performance metrics related to resource efficiency, throughput, and latency. The rapid evolution of large language models (LLMs) has ushered in a [...] Read More... The post Scaling L
Official Groq (LPUs) technical update and publication covering Groq Partners with Aljammaz Technologies to Power AI Inference Across MENA.
Official technical announcement and publication from Hugging Face covering AI for Food Allergies.
General Science
We trained SWE-grep and SWE-grep-mini, fast agentic models specialized in highly parallel context retrieval. They match the retrieval capabilities of frontier coding models, while taking an order of magnitude less time. Available now in Windsurf’s new Fast Context subagent, and our new SWE-grep demo playground!
Official technical announcement and publication from Hugging Face covering Google Cloud C4 Brings a 70% TCO improvement on GPT OSS with Intel and Hugging Face.
October 16, 2025
6584 articles sourced historically · 100 per page