AI NEWS DAILY
Hardware
Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
DeepSeek DSpark in llama.cpp: How to Get 2x Local Inference on…
MacPaw taps Liquid AI to offer on-device inference to devs building for…
Breaking | DeepSeek V4 Flash 0731
Breaking | Oracle bans AI-generated code from OpenJDK
Breaking | Databricks drove down AI coding spend 70%
AI psychosis is the new leadership blind spot
Alibaba tests new business model for Qwen open-source AI
Kitesurf: Agent-first browser that runs in V8 isolates
Cloudflare launches Kitesurf, a browser built for AI agents
AMD acquires Taalas to boost inference performance by etching models in…
I won't read LLM authored fiction
Anthropic CEO reportedly worried new hires only care about money
Qwen3.8 Max now ranked as the best overall model by agentic index
New Orleans is testing Carbyne’s AI-powered Emergency Call Triage…
Software development with AI is starting to feel like cooking steak
xAI, SpaceX, and the Race for AI Buildout
Humans missed 1 in 3 threats approving AI agent commands across 40k…
deepseek-ai/DeepSeek-V4-Flash-0731 · Hugging Face
Video | The OpenAI–Hugging Face Incident [video]
DeepSeek planning to significantly raise prices
OpenAI improves GPT-5.6 Sol in ChatGPT and restricts free users to its…
Microsoft filings suggest "around 70%" of its AI revenue is on OpenAI
Born Against, or why hobby programming communities are against LLM usage
Meta Ran Ads That Contained AI-Generated Child Sexual Abuse Imagery
Launch HN: HyperProbe (YC S26) – Agents that do read-only debugging in…
Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
Prime Agent: A self-improving RLM agent
LLMs won't break symmetric crypto
Cloudflare OS: an open platform for agents, apps, and work
Governments are making a dangerous bet on the AI boom
Mistral's open model Shieldstral matches much larger safety models at a…
Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence…
When online commenters detect my art as AI
I’m leaving OpenAI to build telepathy
Position: LLMs Can't Jump
TIME Is Serving AI Bots a Different Website, with Ads Built In
OpenAI says my prepaid credits were consumed, refuses to show any record
Why Erdős Problems Are Falling to AI
After Rippling blew millions on AI in months, it built an employee ROI…
Building an Advanced Agentic Harness
Microsoft's AI Sales Mostly Come from OpenAI, Disclosures Show
OpenAI flags its new Astra model as potentially reaching the highest…
Iowa-led states ask OpenAI to keep their bots on a leash
AI music generator Suno tightens rules to fight spam and address…
Meta launches Muse Code, an AI agent for large code bases
Awareness – local-first AI agent memory, 96% R@5 on LongMemEval…
AMD acquires Taalas, a startup that bakes AI models directly into…
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
Anthropic loosens Fable 5's biology restrictions but keeps the…
OpenAI's first smart speaker is expected in 2027 at over $300
Stanford Evo 2 AI model generates phages against E. coli
Responding to the next frontier of critical cyber capabilities
Airbnb says AI is helping it ship features faster as it tests a new…
Jill Lepore on the ‘Artificial State’ and why Silicon Valley’s leaders…
OpenAI's hockey-puck-sized smart speaker with moving parts is set to…
Rust-lang/rust is adopting an LLM policy
How AI Is changing Instagram engagement without replacing the human…
China's Largest AI Model Is Being Developed at Bytedance
Stanford and Arc Institute scientists used AI to design new viruses…
AI fuels more than half of cybercrime in Africa as scams surge –…
New Mexico court orders Meta to pay additional $567M in child safety…
Zero-Mem: Zero-Token Memory Operations for LLM Agents
Amazon, Cursor, Microsoft, OpenAI, and Vercel unite on a shared…
Alibaba, DeepSeek push China’s AI model race towards lower costs
Apple says more ex-employees may have taken confidential data to OpenAI
How HSP GRUPPE builds AI capabilities for tax advisory
AI-Generated Images Discourage Me from Reading Your Blog
Security Incident INC-2026-07-28-01 – UK AI Security Institute
I Made My Evals Replay Every Task on a Local Model. The Frontier Lead…
The AI Demand Bubble
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
The Warp Agent CLI
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure…
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in…
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic…
It's not a fear of "AI communism"; it's a fear of competitive market…
TS-RAG: Retrieval Augmented Generation for Time Series Forecasting
AI Data Centers Are Driving Up Power Bills – This Map Shows Where
OpenAI’s new AI smart speaker will reportedly sell for between $300 and…
Scalable estimation of VARMA models
BaKron: Efficient Quantization with Kronecker-Factored Hessians
Stochastic Dynamics on Persistence Diagram Space via Reinforcement…
Surv-IPTB: An Attention-Based Model for Estimating Individual…
Microsoft's AI revenue reportedly depends on OpenAI for 70 percent
From Passive Mirrors to Active Agents: Holonic Digital Twins for…
NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering
MetaboLLM: a metabolomics-specialized large language model for…
Benchmarking and Enhancing LLMs for Rule-Intensive Review of National…
Bias Analysis of L2 Speaking Assessment Systems Using Concept…
QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI…
EnvACE: Internalizing Environment Dynamics via World Rehearsal for…
Deepmind's talent drain likely comes down to chip shortages, a conflict…
Why Large Language Models Fail at Tabular Prediction
Claude Code is the fastest agent framework but costs nearly three times…
ChatGPT brings unlimited text chats to free users
Naïve raises $28.5M to automate the grunt work of setting up and…
The company that made open weights mainstream now competes on discounts
Gen Z dating apps like Ditto ditch swiping in favor of AI matchmaking
Why health AI interfaces must adapt to user expertise
Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for…
OpenAI says Apple’s own security practices undermine its trade secrets…
OpenAI reportedly slows research after its own models secretly…
Amid legal battles, Suno says it will start watermarking songs
OpenAI developer warns the "tireless eagle eyes of a million models"…
Ex-Spotify employees raise $10M to bring the AI behind its…
Exclusive: Mirendil inks $100M+ Google Cloud deal to scale…
WeatherNext: AI model achieves breakthrough in forecasting cyclones
Omilia raises $67M to scale its customer support platform
Google Maps adds agentic features, including food ordering and hotel…
Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna…
Working with the American Psychological Association on youth mental…
From asking to doing: How the world is putting ChatGPT to work
Google Deepmind loses both its CEO and chief scientist as Demis…
Google will shut down Google Assistant starting September 2026 as…
Klaviyo acquires Elias Torres’ Agency in full-circle reunion for tech…
UK's job market is splitting in two as AI demand surges while knowledge…
Jeff Dean and other top AI researchers are leaving Google to launch…
SpaceX’s ambitious compute goals could require over two million Nvidia…
Black Forest Labs makes FLUX 3 Video generally available and claims it…
US appeals court allows Perplexity's AI shopping agent back on Amazon
An AI agent went rogue during UK safety tests, creating fake identities…
Hark previews its browser use agent for completing tasks
Anthropic is hiring an AI chip design team
Shopify says AI search is driving more traffic and sales, not replacing…
TechCrunch Disrupt 2026’s Real World AI Stage features robots…
PRISM2 model uses clinical dialogue to interpret pathology slides
AI makes weather prediction better. Can WindBorne make it lucrative?
SpaceX has bought $329M worth of Tesla Megapacks so far this year
Nvidia doesn’t mess around: A week after open AI industry group formed…
Anthropic signs $10B deal with AI cloud startup Volta
This year's Pulitzer Prizes saw a record number of winners disclose AI…
Open-weight AI models are catching up to the frontier. The safety gap…
Meet Wrinkles, an app that uncovers the hidden stories of the places…
Texas halts new data centers as governor calls for audits
Elon Musk spends half his time talking robots and AI on Tesla earnings…
Google moves billions in Anthropic chip risk off its balance sheet
Anthropic locks in $10 billion of compute from Volta, a cloud startup…
Silicon Valley’s rift over open source pushes back contemplated White…
Spotify expands AI remix and covers project with Merlin partnership
Third-party cyber evaluations involving OpenAI models
OpenAI fires back at Apple's trade secret lawsuit with chat logs…
EON wants to move the data superhighway from ocean fiber to space lasers
Is the future of data centers portable? Runware builds a pod to find out
Red Hat, NVIDIA, IBM back project turning AI policy into code
Research Papers (135 entries)
RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for…
Tytan: Interactive Neurosymbolic Construction of Analytic Semantic…
Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational…
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
DASH: Divergence-Adaptive Supervision Horizons for On-Policy…
A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with…
On-Policy Self-Distillation without Any Supervision
The Bitter Lesson of Tool Calling
The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
CalibForge: Adversarial Solver Calibration for Scaling Learnable…
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning…
RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for…
What Current AI Benchmarks Leave Unmeasured: Modality, Search…
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid…
OTLesMix: Wasserstein Barycenter and Optimal Transport Map for…
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
The Low Frequency Trap: Video Language Models Fail at Simple Event…
Does FLAIR super-resolution erase or hallucinate small white-matter…
Improving the Realism of Synthetic Clinical Benchmarks Under Utility…
Toward Deployable Bangla Sign Language Recognition with…
Depth-Guided Video Object Counting in Crowded Scenes
Learning When to Trust via Selective Context Preference Optimization
Resourced Authority A Mechanism-Design Model for Participatory…
Timestep-Conditioned Transformers for Global Weather Forecasting
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image…
Investigating Artificial Intelligence Digital Sovereignty in Mobile…
Robot Learning from Human Demonstrations: Handwritten Alphabet…
Hypothesis Testing with Conditional Queries: Learnability and the Value…
Muon on the Stiefel Manifold Admits an Exact Closed-Form Update
An Optimal Agnostic PAC Algorithm
Beyond Marginal Validity: Finite-Sample Guarantees for Localized…
Continual Learning in Transition
Optimal Rates for Learning with Monotone Adversaries
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture…
Challenges in Evaluating Explanation Methods for Static and Evolving…
Hierarchical Graph Memory for LLM Agents with Path-level Localization…
Chained Recursive Language Models for Multi-Iteration Reasoning
DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model…
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances…
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of…
BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery…
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
Spoken Function Calling: A New Perspective on Spoken Language…
Stable Density Ridges: Consistency and Convergence of Subspace…
MultiPathFormer: Towards a Foundation Model for Multipath Wireless…
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training…
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality…
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked…
MarsCast: Transfer Learning of AI Weather Foundation Models to…
SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns…
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Optimizing What Policies Learn From: Recoverability-aware Rollout…
The Effect of Perceived Race and Gender on Police Language Use…
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Reasoning Core: Designing Broad Procedural Data for…
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and…
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom…
Hardware Design and Security in the Era of Chiplets and LLMs
From Score Matrices to Football-Aware Match-State Simulation: An…
ORACLE: A Multi-Objective Reinforcement Learning-Based Analog Circuit…
Protoreasoning in Tiny Transformers
MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning
Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for…
Link prediction on multi-relational graphs from an influence…
Reward Structure Shapes the Interaction Between Episodic Exploration…
Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous…
Multimodal Spatiotemporal Atmospheric Data Assimilation with Latent…
ArtAnno: Annotating Implicit Semantics in Artworks through LLM…
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic…
CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning…
German parties shifted towards intuition-based rhetoric after the far…
VQ-VAD: Vector-quantized Motion Representation Learning for…
Short-term load forecasting under EU-AI Act Requirements in…
Item Response Theory for AI Safety
Stochastic Emulation using Generalized Stratified Sampling for…
Provable Limits and Certified Deferral for Verbalized Uncertainty in…
RepairFormer: Automated Repair of Structured Inputs Using Transformers
The Loss Does Not See the Basis, but Adam Does
Representational separation between unitary and channel quantum…
Same Formulas, Different Semantics: Do Language Models Follow Modal…
Language Models Generalize to Human-like Word Order Preferences
Revealed Rationality: Label-Free Evaluation and Regularization from…
Canonical Joint Energy-Based Model on CIFAR-10: failure modes and…
SciCode-Verified: How Benchmark Defects Underestimated the…
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and…
ReflectRL: Learning from Golden Negative Trajectories via…
When and Where to Look: Adaptive Visual Evidence Scheduling for…
ParVL: Parallel Scaling and Expandable Compute Allocation for…
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch…
Information-Geometric Forward Policy Training in GFlowNets
Equivariant Music Transformer
Cross-Model KV Cache Transfer in LLM Families: A Closed-Form Linear…
GENESIS: Towards Explainable Causal Discovery
Interpretable Adaptive Sampling for LLM Test-Time Scaling
Trajectory inference via Acceleration Matching
Sparse Weight Decomposition for Efficient Circuit Extraction
MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for…
A game theory for foundation models shows new paths to rational…
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs…
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated…
When Attention Goes Blind: Numerical Failure in ALiBi Positional…
Assessment of Conditional Diffusion Model for Synthetic Histopathology…
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English…
Latent Reward Registers for Diffusion Preference Alignment
Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
SocietyBench: Forecasting Counterfactual Social-World Evolution
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement…
Can Large Language Models Recover Semantic Optimization Opportunities…
Separating quantum circuits from classical LLMs
PRISM: Powerful Time Series to Image (TS2I) Representations for…
When Efficiency Becomes Fragility: Exploiting Dynamic Routing…
ANNOTARES: A Dataset for Extracting Logical Structures from German…
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary…
BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for…
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Beyond Representational Similarity: Source-Conditioned…
Agogic: Performance-Timed Music Tokens for LLM-Native…
A Physics-Flavored Transformer Network for Parametrizing Contraction…
The Transformer Revolution, Part 1: Dynamic Processing through Output…
Socially Grounded Agentic AI: Coordinating Plural Perspectives through…
Omega-S: A Functional Resilience Index for LLM Fine-Tuning
DS@GT-ARC at eRisk 2026 Task 3: Sparse, Semantic, and LLM Reranking for…
SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for…
Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice…
HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for…
Logic Before Language: Pre-pretraining on Formal Derivations Fosters…
Implementing Causal Perception: Competing SCMs and Situated Fairness
CRS-Triage: Confidence- and Reliability-Aware Selective Triage under…
string2string Studio: An Interactive, In-Browser Platform for…
Muon Meets Mamba: Spectral Optimization for State Space Models
Operationally Feasible Synthetic Power-Grid Scenarios via Learning the…
ADMITBench: A Safety-Governed Reference Framework for Evaluating the…
Robust Low-Tubal-Rank Tensor Completion under Cross-Concentrated…
Intertemporal Preference Steering in Qwen3 via Contrastive Activation…
Archive
New Framework Desktop Option with AMD Ryzen AI Max+ Pro 495 and 192GB…
Nvidia unveils RTX Spark Superchip — Arm CPU + Blackwell GPU + 128GB…
GPU drought 2026: no new gaming GPUs from Nvidia or AMD this year
Same Score, Different Brains: Coding Agents on an RTX 4090 — wasnotwas
Anthropic will deploy 2 gigawatts of AMD GPUs for Claude in a deal…
SkewAdam: A tiered optimizer that cuts MoE state memory by 97% (fits a…
The AI Slot Machine Effect: Why Generative Feeds Disrupt Deep Work And…
SpecJudge: a local-first CLI that reads your project specs and tells…
Reproducing OpenAI’s “persistently beneficial models” - GRPO trait…
llama-bench MiMo-V2.5 UD-Q6_K_XL on 2x Strix Halo with thunderbolt + 2x…
Help me evaluate a 4-layer Al homelab architecture
Benchmarked every spec-decode method on Qwen3.6-27B across vLLM and…
Those in the 1000+ prefill and 100+ decode range on Qwen3.6 35B at Q4…
I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on…
Dual V100 SXM with NVLink: which PCIe configuration?
Unsloth now supports AMD!
543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a…
How to Train a Gen AI Kick Drum Model on Your Old Linux Desktop with…
How To Get Webdesign Projects?
LLM Networking with MikroTik
Bonsai 27B is a full open reasoning model that fits on an iPhone
Developers Hate AI. I Used It To Sell 10 Websites This Week.
Deepmind CEO Hassabis says "nobody in the world knows what happens…
Upgrade path for ryzen 9 (64 gb) + rtx 5080
I benchmarked 15 "E-Waste" GPUs with Modern Workloads
AI agent crawlers now need permission. Here’s how to get it
Prompt-engineering paper accepted to ICML [R]
Running Qwen3.5-122B on Mac Studio 96GB: Fixed 3 bugs that made…
Ph.D. in Operations Research / Big Tech Eng: How to transition into…
Local Image to 3D (<2gb RAM, <20s, Apple Silicon, iPhone)
I got Nemotron Puzzle 75B running smoothly on a 64GB M2 Max
Benchmark - 4x 5060 Ti (64GB VRAM) (P2P) - Qwen3.6 27B (INT8 /w bf16 kv…
i would like to share my experience. working with huge LLMs and poor…
**Your $80 Tesla P100 has been doing silently noisy math in llama.cpp…
First attempts at a CPU setup - MS-02 Intel 285hx, trying Qwen3…
Qwenthropic
Measuring PCIe transfer under dual GPU with pipeline & tensor llama.cpp
Ultra budget 20GB vram with 448GB/s for $100 bucks.
I benched quad 5060Tis for code generation with Qwen3.6-27B so you…
Ask HN: How do you use Vim in the era of AI?