Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execut
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolu
343 points, 416 comments
The post When AI Stops Taking Notes and Starts Seeing Patients appeared first on Lightspeed Venture Partners .
Official technical announcement and publication from Hugging Face covering Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis.
1041 points, 452 comments
Learn how AI cost management helps organizations move from AI pilots to measurable ROI through greater visibility, governance, and optimization. The post The Economics of Agent Optimization: From pilots to measurable returns appeared first on Microsoft Azure Blog .
A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research .
632 points, 615 comments
304 points, 228 comments
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
With Sierra’s long-running Horizon agents, insurers can now proactively engage every prospective customer over days, weeks, or months — staying with them until they buy or move on.
1009 points, 936 comments
Affine frequency division multiplexing (AFDM) is a promising chirp-based multicarrier waveform for high-mobility integrated sensing and communication (ISAC). Accurate angle, delay, and Doppler estimation is essential for AFDM sensing. Since target delays and Doppler shifts are generally continuous-valued, representing them on a discrete delay--Doppler grid causes energy leakage and peak displaceme
Official CMU (Carnegie Mellon AI) technical update and publication covering Shen Discusses AI Safety at WEF Annual Meeting.
In this work, we construct fast eighth-order Pade schemes for the direct Zakharov-Shabat scattering problem. The schemes are based on an eighth-order exponential integrator obtained from the Magnus expansion. A direct extension of the conventional fast Pade representation to the eighth-order case leads to insufficiently accurate fast variants in the considered tests. To overcome this difficulty, w
Generative AI
We investigate how small a co-propagating transverse-scalar component can be resolved by Taiji in an already identified bright tensor chirp. Using a source-tracked tensor-null response, we formulate the problem in terms of the minimum resolvable scalar strain fraction and evaluate it over the sky. For a one-year benchmark chirp with tensor signal-to-noise ratio $ρ_T=1000$, we find that Taiji can r
OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption.
Official Groq (LPUs) technical update and publication covering Groq Becomes an NVIDIA Cloud Partner.
See how RingCentral uses ChatGPT Work and Codex to accelerate AI product development and centralize operational intelligence across engineering and operations.
Official Harvey AI technical update and publication covering The Legal Team's Guide to Employment Contract Drafting.
Introducing Grok 4.6
Official Fireworks AI technical update and publication covering Can open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on Kimi K3 & Qwen3.5-9B.
Official Artificial Analysis technical update and publication covering Upstage Solar Pro 4: Benchmarks and analysis.
Official Deepgram technical update and publication covering Podcast Transcription.
Official Artificial Analysis technical update and publication covering Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency.
Official LMSYS Chatbot Arena release and benchmark update covering SGLang and Miles Add Day-0 Support for Qwen3.8.
Official Decagon technical update and publication covering Getting closer to the build: How I helped create agent deployment engineering.
Official Harvey AI release and benchmark update covering A Smarter Inbox Built for Legal Work: The New Harvey for Outlook.
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framewo
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers grounded in live web results. We developed OreoLook (formerly lixSearch), an open-source answer engine using automated browser agents and provider-routed LLM inference. Its local search, caching, session-management, and embedding stack runs on commodity CPU hardware; answer synt
On 2025 August 18, the LIGO-Virgo-KAGRA collaboration reported a sub-threshold gravitational-wave (GW) candidate, S250818k, consistent with a binary neutron star (NS) merger potentially involving a sub-solar-mass compact object. Follow-up electromagnetic (EM) observations identified a Type IIb supernova, SN 2025ulz within the broad localization area of the GW signal. This potential link between a
Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on dual-encoder models that align visual and textual representations through contrastive learning. The March 2026 release of Gemini Embedding 2, Google's first natively multimodal embedding model to map text, images, video, audio, and documents into a single shar
350 points, 334 comments
Health & Bioscience
440 points, 543 comments
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research .
<div class="block-paragraph_advanced"><p><span style="vertical-align: baseline;">For organizations deploying AI agents at scale, there’s often a critical divide between structured and unstructured data. While large language models (LLMs) excel at parsing text documents, emails, and PDFs, they can struggle when presented with raw enterprise databases. Meanwhile, standard natural-language-to-SQL (NL
Agents built on Sierra can now navigate the most challenging IVR systems out of the box: they understand when to call, how to get through a multi-level menu, when to stay silent, and how to recognize when they’ve reached a person.
Official technical announcement and publication from Hugging Face covering Thinking of ACE? We Can Do It with Fewer Tokens.
524 points, 489 comments
In-region inference, open models, and new European infrastructure for sovereign AI.
Professor Abate will support the AI Office’s scientific assessment of general-purpose AI models, alongside its work on AI innovation, testing and model evaluation.
OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.
OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
# Abstract Thermodynamic concepts are increasingly used in nonlinear photonics to describe Rayleigh--Jeans thermalization, optical wave turbulence, condensation, negative-temperature states, and statistical mode locking. This raises a broader question: how far can thermodynamic reasoning be extended to localized structures maintained far from equilibrium by a balance of gain, loss, dispersion, and
Official Harvey AI technical update and publication covering Enterprise Legal Management and the Spend Lever it Can’t Reach.
Soniox TTS v2 brings extraordinary voice quality, expressive control, exceptional precision, high-quality voice cloning, and low-latency streaming in 60+ languages.
Official Harvey AI technical update and publication covering How to Choose Legal Document Review Software in 2026.
[ Product ]
Official Harvey AI technical update and publication covering Regulatory Change Management: What Happens After the Alert.
938 points, 988 comments
Soniox TTS v2 brings extraordinary voice quality, expressive control, exceptional precision, high-quality voice cloning, and low-latency streaming in 60+ languages.
Official LMSYS Chatbot Arena release and benchmark update covering Unified Radix Cache: One Tree for Hybrid Model Prefix Caching.
Official Mistral AI technical update and publication covering In-region inference, open models, and new European infrastructure for sovereign AI..
Official LMSYS Chatbot Arena release and benchmark update covering SGLang Adds Day-0 Support for NVIDIA Nemotron 3.5 Lightning.
Official Stanford (HAI) technical update and publication covering Companies That Buy and Sell Your Data Are Not Following California’s Strict Privacy Laws.
451 points, 425 comments
The space-based detector Laser Interferometer Space Antenna (LISA) will observe inspiralling black hole binaries in the mHz band, many of which may retain orbital eccentricity. In contrast to quasicircular binaries, eccentric systems radiate through multiple orbital harmonics, while relativistic periastron precession introduces a secular phase structure in the waveform. We exploit this structure t
“GeoPT” helps AI models understand the basics of physics so they can simulate how objects respond to things like wind and water more efficiently and accurately.
534 points, 185 comments
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated fea
OpenAI CFO Sarah Friar shares five lessons for building an AI-native finance function, from automated forecasting to stronger controls and AI ROI.
Official technical announcement and publication from Hugging Face covering Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS.
<div class="block-paragraph_advanced"><p><span style="vertical-align: baseline;">At Google Cloud, we help organizations of all sizes build and operationalize complex agentic workflows with total confidence. By combining world-class AI research with an open, fully integrated AI platform, we give customers the flexibility to innovate and the foundation to deliver measurable business value. </span></
<div class="block-paragraph_advanced"><p><span style="vertical-align: baseline;">What’s the best way to recommend products to little-known users? </span></p> <p><span style="vertical-align: baseline;">We’ve spent our careers trying to solve this problem for major companies like Spotify and Priceline, and it’s why Sidd founded </span><a href="https://www.malachyte.com/" rel="noopener" target="_blan
<div class="block-paragraph_advanced"><p><span style="vertical-align: baseline;">Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and optimize their ad spend. WPP is replacing that guesswork with an AI-powered view of shifting market dynamics, giving bran
Femtosecond lasers underpin applications ranging from material processing to corneal surgery, while their regular pulse trains form optical frequency combs that have revolutionized timekeeping, spectroscopy, and metrology. On-chip optical frequency combs, such as Kerr microcombs, have enabled high-repetition-rate applications in optical communications and microwave photonics. However, integrated c
642 points, 601 comments
OpenAI sent Governor Greg Abbott a letter outlining its commitment to responsible AI infrastructure in Texas. The letter supports reliable, transparent growth that benefits Texans.
Official ElevenLabs technical update and publication covering Customer Stories.
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
Official technical announcement and publication from Oxford (AIDC) covering Research Associate on ERC Synergy VePaSS.
Official technical announcement and publication from Oxford (AIDC) covering Research Assistant on AI Safety.
The award is given to a promising young researcher working in the area of constraint programming.
1208 points, 639 comments
Official technical announcement and publication from Hugging Face covering Making Knowledge Distillation Cheap Enough to Run at Scale.
Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit validation, and security testing.
Approved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, governed cybersecurity services to customers.
We present timing and spectral analysis of the recently identified ultracompact double-degenerate (DD) white dwarf binary eRASSU J060839.5$-$704014 using observations from NICER and Einstein Probe (EP), together with archival XMM-Newton data. By phase-connecting the long-term XMM-Newton, NICER, and EP observations, we obtain a coherent quadratic timing solution, yielding an orbital period of 374.1
694 points, 396 comments
Virgin Atlantic is accelerating research, product planning, and decision-making with ChatGPT Work, helping teams connect signals across the customer journey.
The enterprise marketing team at Zapier uses ChatGPT Work to reduce the number of drop-offs in its lead funnel, build campaign assets, and automate reporting.
Official Harvey AI technical update and publication covering Practice-Ready From Day One: How SMU Dedman Teaches Legal Writing With Harvey.
Official Harvey AI technical update and publication covering How In-House Compliance Teams use AI to Stay Ahead of Regulatory Change.
Premium seats are coming to ChatGPT Business. Sign up by August 20 to get $100 in workspace credits and unlock higher usage for your team's most demanding work.
Official technical announcement and publication from Hugging Face covering Meta is back with Muse Glimmer: local, agentic, multimodal, and open source.
Official Stanford (HAI) technical update and publication covering See All Media Mentions.
Official Stanford (HAI) technical update and publication covering New Stanford Grants Tackle AI's Impact on Global Security and Geopolitics.
Official Fireworks AI technical update and publication covering Muse Glimmer from Meta on Fireworks: Ideal for your Always-On Agents.
Official LMSYS Chatbot Arena release and benchmark update covering SGLang Adds Day-0 Support for Muse Glimmer, a Multimodal Model Built for Local Agentic Workflows.
438 points, 380 comments
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and edi
Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harness---is typically treated as a fixed artifact after deployment. This work studies an alternative where the harness is task-specific and continuously evolvable: each task family maintains its own harness, which is hot-swapped across iterations throu
449 points, 130 comments
Aug 8, 2026
The post Phil Duggan appeared first on Lightspeed Venture Partners .
316 points, 268 comments
6542 articles sourced historically · 100 per page