Install
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation
2+ day, 3+ hour ago (279+ words) Is it deployable? Yes. The code is Apache-2.0, the ToolGrad-500 dataset and the 1B, 4B, and 12B models are on Hugging Face, and there is a PyPI package. ToolGrad reverses the order. It first constructs a ground-truth tool-use chain by actually executing APIs,…...
d-Matrix and NVIDIA Plan NVLink Fusion Rack System for AI Inference
2+ day, 13+ hour ago (531+ words) “Purpose-built connectivity is what turns innovative compute into high-performing AI factories,” said Jitendra Mohan, CEO of Astera Labs. “Our partnership with d-Matrix and NVIDIA brings this vision to life within the NVLink Fusion ecosystem, delivering high-throughput for low latency AI…...
d-Matrix Adopts NVIDIA NVLink Fusion Rackscale Infrastructure for Ultra-Low Latency AI Inference
2+ day, 17+ hour ago (347+ words) d-Matrix will integrate XPUs into NVIDIA MGX rackscale architecture with NVIDIA NVLink Fusion; Collaboration to enable AI labs, hyperscalers, neoclouds to offer premium-level token services “Purpose-built connectivity is what turns innovative compute into high-performing AI factories,” said Jitendra Mohan, CEO…...
D-Matrix adopts NVLink Fusion for rack-scale AI inference
2+ day, 18+ hour ago (210+ words) The AI chip startup will plug its next-gen Raptor XPUs into NVIDIA's rack-scale fabric, targeting ultra-low-latency workloads by late 2027 Building your own AI chip is hard. Getting that chip into production racks at scale is arguably harder. D-Matrix, the inference-focused…...
CoreWeave Launches Physical AI Field Engineering
2+ day, 21+ hour ago (298+ words) CoreWeave recognized as a Visionary in the Gartner® Magic Quadrant™ for Cloud AI Infrastructure. Read the report Our AI-native platform of technology, tools, and teams is fully integrated and purpose-built to power pioneers' most complex workloads. The only AI cloud…...
Lightbits launches KV cache engine to boost GPU performance
3+ day, 6+ hour ago (303+ words) Lightbits Labs has launched its AI inference software engine, Infera, aimed at improving inference performance and cost efficiency. The company says Infera expands KV cache data beyond limited high-bandwidth memory attached to GPUs and targets large language model operations with…...
CoreWeave, Parallel Works Support DARPA NODES
3+ day, 18+ hour ago (396+ words) CoreWeave recognized as a Visionary in the Gartner® Magic Quadrant™ for Cloud AI Infrastructure. Read the report Our AI-native platform of technology, tools, and teams is fully integrated and purpose-built to power pioneers' most complex workloads. The only AI cloud…...
Palantir And Nebius Partner To Deliver Sovereign AI Stack To Commercial Customers
3+ day, 18+ hour ago (484+ words) Palantir Technologies and Nebius Group have formed a strategic partnership to bring Nebius’ AI-native compute infrastructure and cloud platform to Palantir’s commercial customers, creating a more integrated sovereign AI offering for organizations that want greater control over their compute, data,…...
Mercury 2.5 vs Nex-N2.5-Pro - AI Model Comparison
3+ day, 21+ hour ago (81+ words) OpenRouter Mercury 2.5 vs Nex-N2.5-Pro: side-by-side summary Mercury 2.5 and Nex-N2.5-Pro are available through the OpenRouter API, so switching between them takes a model slug change rather than a new integration. Mercury 2.5, from inception, has a 260,000-token context window and…...
NVIDIA CUDA applications rely on data transfer between memories
2+ week, 6+ day ago (19+ words) ServeTheHome NVIDIA CUDA applications rely on data transfer between memories...