AI News Archive: August 10, 2026 — Part 24
Sourced from 500+ daily AI sources, scored by relevance.
- A Mechanistic Diagnostic of Rank Collapse in Post-Norm Decoder Transformers
Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes. Although prior work has identified rank collapse and gradient vanishing as rel...
- Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation
Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that make...
- Beyond Binary: Continuous State Optimization with Graph-Structured Objectives
Large-scale learning systems often face the challenge of balancing multiple, potentially competing objectives, such as fairness, accuracy, and latency. While recent work has formalized this as an optimization problem over binary states, many real-world control parameters, such as fairness th...
- In-Context Density Estimation for Tabular Data
Density estimation underlies many unsupervised tasks on tabular data such as anomaly detection, out-of-distribution detection, and data augmentation. Although all these problems reduce to questions about where probability mass lies, they are typically solved individually by fitting a separate model ...
- RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera
Robotic manipulation requires perception systemsthat identify actionable parts such as handles, rims, triggers,and tool tips, not merely object categories or point clouds. This paper presents RoboSeg, a part-level semantic reconstructionsystem that links vision-language model (VLM) functional-partdi...
- SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations...
- Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition
Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, avoiding the sim-to-real gap and enabling continuous policy refinement through human-in-the-loop interaction. Recent methods have demonstrated sample...
- Beyond the Plane: Coupling Planar Vehicle Dynamics with Three-Dimensional Road Geometry
Simulation is crucial for developing and testing autonomous driving systems. In particular, the development of localization and control algorithms relies on an accurate vehicle dynamics simulation. However, most vehicle dynamics models are two-dimensional while real-world roads are three-dimensional...
- DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving
Large-scale language models for autonomous driving enable enhanced global understanding and long-horizon planning. However, when deployed in isolated vehicles, limited sensing range and occlusions restrict reliable decision-making, and the substantial computational and latency overhead makes on-boar...
- Task-Oriented Formation Decision via Reinforcement Learning: Herding an Attacking Swarm
Multi-robot systems can accomplish tasks that are difficult for a single robot by organizing into task-specific formations. Different from existing studies on multi-robot shape formation, we here study the task-oriented formation decision problem, with a focus on the herding task. This task is chall...
- SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning
While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by suboptimal execution speeds. Imitation learning policies are inherently limited by hardware constraints and the speed of the operator during data collection. In addit...
- UnsDrive: Towards Robust End-to-End Autonomous Driving in Unstructured Scenes
End-to-end planning has shown strong promise for autonomous driving, but most existing methods are designed for structured urban roads and generalize poorly to unstructured mining environments. In such settings, weak road structure, terrain-induced occlusions, degraded visibility, and large unobserv...
- HarnessWAM: Bridging Prediction and Deliberation in World Action Models
World Action Models (WAMs) jointly learn environmental dynamics and robot actions, introducing priors over physical evolution into embodied control. However, finite-horizon prediction and action generation are insufficient for complex embodied tasks that require global planning, cross-stage state ma...
- Rethink Before You Execute: Adaptive Execution for World Action Models
World Action Models (WAMs) jointly predict future actions and the evolution of the environment. At each inference, a WAM generates a chunk of actions and the robot executes a fixed prefix before replanning. We argue that this fixed execution horizon is poorly matched to execution dynamics: the chunk...
- Graph-Guided Safe Diffuser: Topological Graph Guidance for Safe Diffusion Planning
Many diffusion-based planners enforce safety through inference-time guidance, but such interleaved trajectory deformations often degrade kinematic feasibility due to manifold rupture. We propose Graph-Guided Safe Diffuser (G2SD), a hierarchical framework that leverages a high-level topological graph...
- Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation
Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length history. However, real-world manipulation is often non-Markovian, requiring robots to retain and reason over task-releva...
- JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from t...
- SAFE-CHEM: Uncertainty-Aware Policy Switching for Robust Robotic Chemistry
The deployment of autonomous robotic systems in chemistry laboratories is accelerating experimental workflows and providing the foundational data for AI-driven scientific discovery. However, despite the success of data-driven methods in acquiring dexterous skills, safety remains a primary barrier to...
- WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation, and data generation. Realizing this promise requires precise action-conditioned transitions rather than merely plausible outputs. Yet their applicability remain...
- SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot
Most existing vision-language navigation tasks assume that instructions are complete and unambiguous. However, real-world robots often encounter natural human instructions that are ambiguous, underspecified, or incomplete, requiring them to resolve such uncertainties through active questioning. Inte...
- Ventilate · At the right time
Beat the heatwave: when to open, close, and use a fan
- Particle-Based Conformal Prediction for Contact-Aware Uncertainty Calibration in Stratified Configuration Spaces
Reliable uncertainty representation is essential for deploying autonomous systems that interact with their environment, as robots must reason about how uncertainty arising from both stochasticity and model mismatch is impacted by contacts with obstacles (e.g., when navigating through a cluttered env...
- High Fidelity Capture, Reconstruction, and Transfer of Human Demonstrations for Robot-Assisted Bathing
Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for complex, sustained physical interaction with humans. A key challenge hindering the development of such systems is that collecting, understanding, and effec...
- Trajectory Divergence Horizon Decision for Reliable Dual-Arm Surgical Subtask Manipulation
Surgical robotic systems are increasingly being adopted as clinical workload rises, motivating autonomous solutions for repetitive manipulation subtasks. Learning-based controllers improve generalization compared with rule-based and analytic approaches, but most are trained for individual tasks and ...
- ROEVO: Robust Organized Edge Feature-based Visual Odometry Using RGB-D Cameras
This work presents a visual odometry (VO) system that leverages image edge features. Edges are spatially expressive cues commonly present across diverse environments, offering rich textural and structural information. However, existing edge-based VO methods often fail to fully exploit this potential...
- Latent World Models with Monotone Planning Costs for Image-Goal Navigation
Image-goal navigation with latent world models requires not only accurate future prediction, but also a planning cost that reliably ranks candidate action sequences. We define the cost as the cosine distance between the predicted future embedding and the goal embedding, and show that poor cost order...
- Personalized Lower-limb Exoskeleton Assistance via Preference-based Bayesian Optimization
A significant challenge in exoskeleton robotics is the need to dynamically adapt control profiles to individual motion preferences, thereby ensuring both efficient and comfortable assistance. Currently, since user experience can serve as a comprehensive metric for evaluating the effectiveness of ass...
- Personalized Federated Learning via Variance-Aware Nonparametric Empirical Bayes
We develop a new approach to Personalized Federated Learning across heterogeneous clients using Nonparametric Empirical Bayes (NPEB). Leveraging the asymptotic normality of local parameter estimates obtained from Empirical Risk Minimization or M-estimation, our method formulates these estimates as n...
- CPDA: Class-Conditional Path Distribution Alignment for Unsupervised Time-Series Domain Adaptation
Unsupervised time-series domain adaptation (DA) addresses the challenge of transferring a classifier from a labeled source domain to an unlabeled target domain under distribution shifts induced by different users, sensors, devices, acquisition conditions, or temporal dynamics. Existing methods typic...
- Context Is Not Authority: Structured Runtime Governance for Financial Market Agents
Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present SAGE-Fin, a finance-specific authority-handoff contract that makes the proposed effect, not merely its text, the object of runtime control. SAGE-Fin compiles pro...
- oqoqo
Build evals and custom benchmarks for real-world tasks
- Portfolio Lab
AI investing, done responsibly
- Paritok
Spend up to 85% less and run 3× longer coding agent sessions
- SecondBrain Note by GenSpark
A MagSafe AI Recorder That Acts for You
- AI Group Call
Type a goal, join a live voice call with six AI minds
- Prime Agent
A coding agent that can refine its own harness
- Gutta
A tiny, offline task list for your Mac menu bar
- Remix
Figma, but on your production app. Test variants and ship.
- Vidaya
Healthspan score from your wearables, labs, and DNA.
- Salesman AI
The AI sales agent that turns meetings into revenue
- Heym
Build agentic systems. Run them with confidence
- Account Moodboard
Visualize Instagram & TikTok accounts' vibe & stats
- VICE Platform - Private Beta
Security scans for people who ship fast
- FirstSignal
AI voice agent for first-round interview screening
- Aarambha: Learn & Invest
Duolingo X Spotify for Investing
- buildbook
Share Raw, Unfiltered Behind-the-Scene Moments To The World!
- Compartment Enterprise
Securely share and scale custom internal apps and workflows
- GiraffeDoc
White-label API docs sites - hand-built for $45/mo
- OfferTrail
Turn job-search emails into an organised application tracker
- Corteks
Explore your mind