AI News Archive: June 25, 2026 — Part 20
Sourced from 500+ daily AI sources, scored by relevance.
- Associations Between TMS-Induced Electric Fields and Craving and Consumption Outcomes in Substance Use Disorders: A Multimodal Dose-Response Meta-Analysis
Background: Transcranial magnetic stimulation (TMS) is a promising treatment for substance use disorders (SUDs), although heterogeneous stimulation parameters hinder the identification of optimal strategies. Using meta modeling, we linked treatment effect sizes (Hedges' g) to simulated electric field (E field) distributions to identify brain regions associated with efficacy variability. Methods: TMS trials in individuals with SUDs published through the end of 2025 were identified through a systematic PubMed search. Studies reporting craving or consumption outcomes with quantifiable effect sizes were included. Objectives were to (i) examine associations between study-level effect sizes and simulated local E field strength in MNI space for craving and consumption outcomes; (ii) generate a combined E field effect size association map; and (iii) assess spatial overlap with fMRI drug cue reactivity patterns in 60 individuals with SUDs. Results: The analysis included 81 randomized controlled TMS studies, yielding 107 effect size estimates for craving and consumption (n = 75 and n = 32, respectively). Compared with sham stimulation, TMS produced small-to-moderate improvements in both outcomes. E-field modeling identified the pre-supplementary motor area (preSMA) and inferior frontal gyrus (IFG) as regions associated with variability in craving-related effect sizes, and the frontopolar cortex with variability in consumption-related effect sizes. Correlation maps were highly robust (mean leave one out similarity r = 0.996), and the frontopolar cluster showed significant spatial overlap with fMRI drug cue reactivity patterns (Dice coefficient = 0.37). Conclusion: These findings identify frontopolar, preSMA, and IFG regions where local E-field strength is associated with SUD treatment effects, supporting more precise neuromodulation strategies.
- A Reproducible Clinical Decision-Support Suite on MIMIC-IV
Most published clinical-AI results are single models on a single dataset, difficult to reproduce, and rarely validated outside their training hospital. We built a broad, methodologically rigorous, reproducible clinical decision-support (CDS) suite spanning four families - intensive-care deterioration and outcomes, emergency- department triage, electrocardiographic interpretation, and clini- cal natural-language processing - comprising 26 models. Tabular models are gradient-boosted trees over point-in-time, leakage- safe first-24-hour features; deep models include one-dimensional convolutional networks on raw 12-lead ECG, fine-tuned clinical transformers, and an instruction-tuned large language model for discharge-summary drafting. Every model uses patient-level data splits, probability calibration, a shuffled-label leakage gate, and SHAP explanations, and is characterised by its full confusion matrix with sensitivity, specificity and predictive values. Dis- crimination matched or approached published benchmarks: ICU mortality AUROC 0.884, acute kidney injury 0.830, prolonged stay 0.813; emergency-department-to-ICU 0.875; cardiologist- labelled ECG diagnosis 0.909; full-note diagnostic coding 0.892. Raw-signal ECG deep learning improved myocardial-infarction detection by +0.142 AUROC over interval features. The MIMIC- trained mortality model generalised to a different multi-centre US cohort (199,133 stays) with only a 0.044 AUROC drop. We describe how each model family is incorporated into the latest version of the zMed Critical Care application and its CDS tools
- Time-Aware Contrastive Transformer for Longitudinal Patient Representation Learning
Learning high-quality longitudinal patient representations from irregular electronic health records (EHRs) is essential for understanding heterogeneity in time-evolving diseases such as cancer. Longitudinal patient representation learning methods often rely on external labels for downstream tasks or do not model the temporal dynamics between medical events explicitly, reducing the clinical applicability of learned disease trajectories. In this work, we propose the Time-Aware-Contrastive-Transformer (TACT), a transformer-based model that integrates explicit temporal modeling with a fully self-supervised contrastive learning framework. We introduce a sampling-based data augmentation workflow that leverages hierarchical taxonomies of diagnoses and medications to enrich representation learning. Evaluated on a large real-world dataset, TACT demonstrates robust performance across patient representation and event embedding metrics and outperforms two time-aware transformer comparison models. Unlike the comparison models, TACT successfully bridges contrastive learning with medical hierarchies, allowing it to track precise disease trajectories and discover clinically actionable patient phenotypes. Consequently, this approach establishes a comprehensive framework for characterizing patient heterogeneity through the identification of potentially clinically meaningful subgroups with distinct progression profiles.
- A Natural Experiment Reveals Clinically Essential and Compliance-Driven Nursing Documentation
Despite contributing substantially to clinician burnout, nursing documentation lacks empirical evidence distinguishing clinically essential from administratively driven documentation. Exploiting a COVID-19 documentation relaxation policy as a natural experiment, we analyzed 520,357 patient shifts from 36,321 patients in 54 inpatient units (2019 - 2022) using large language model-assisted flowsheet classification and structural equation modeling. When permitted, front-line nurses reliably distinguished two types of documentation: in acute care units, primary nurses reduced compliance-driven Cares & Safety documentation by 19% (106.4 to 86.2 entries, r = -0.19), while maintaining or increasing documentation directly relevant to respiratory management, with no impact on patient respiratory outcomes. Documentation intensity also co-varied with real-time patient deterioration, consistently across unit types (|{beta}| = 0.13 - 0.14). Together, these findings provide the first large-scale quantitative evidence distinguishing clinically essential documentation from compliance-driven documentation and demonstrate that targeted reduction of the latter is a viable strategy for alleviating documentation burden without compromising care quality for respiratory care management.
- Predicting Depression and Anxiety Progression in Multiple Sclerosis from Longitudinal Clinical Data Using Machine Learning
Depression and anxiety are highly prevalent in multiple sclerosis (MS), yet tools for predicting mental health trajectories from clinical data remain limited. We investigated what structured electronic health record data can predict about depression and anxiety progression in MS, and where its limits lie. We developed gradient boosting models to predict PHQ-9 (depression) and GAD-7 (anxiety) score change using EHR data from 2,163 MS patients (7,327 observations) and 1,465 patients (3,319 observations), respectively. Models achieved R^2 of 0.22 (PHQ-9) and 0.28 (GAD-7). Baseline score was the dominant predictor, but this largely reflects regression to the mean: patients with high baseline scores tend to improve, while those with low scores tend to worsen. Age emerged as a consistent secondary predictor across both models: younger patients showed smaller improvements independent of baseline severity. Feature importance differed between models---PHQ-9 prediction relied on symptom subscales while GAD-7 incorporated pain and disease duration. These results suggest that structured clinical data alone capture only a fraction of what drives mental health trajectories, and that richer data sources---clinical notes, patient-reported outcomes, digital phenotyping---will be needed to enable meaningful individual-level prediction.
- Heron
Wireshark for AI Agents: passive eBPF observability
- Milestones
Native project planning app, now on Mac & with an MCP server
- Nashra
Turn followers into clients.
- Paybond CLI
Safe agent spend from the terminal
- Grass 2.0
The always-on computer for your coding agents
- Signspell
Real-time ASL alphabet recognition in py ,pip install and go
- Silica
Your private local AI workspace for Mac
- Seavid AI | Image to Video Generator
Turn ideas into stunning AI videos instantly.
- TopAITools4U
Find the best AI tools for every workflow
- Evaluating Generative Video AI for Standardized Psychiatric Patient Simulation With Graded Hygiene Deterioration.
Abstract Introduction: A clinician's initial assessment during the mental status examination (MSE) places substantial weight on a patient's general appearance, grooming, and hygiene. However, the logistical difficulty of producing simulated or standardized patient (SP) videos that systematically manipulate these characteristics limits the development of clinical AI tools and training curricula. This pilot study investigates the technical feasibility of using a video-generation diffusion model to re-animate modified reference images onto driving videos, enabling the creation of diverse patient presentations without the need for repeated filming. Methods: Utilizing an established publicly available dataset, we extracted reference images of three SPs and applied a text-to-image AI model to generate five appearance conditions: the unmodified baseline and four escalating hygiene-deterioration levels: mild, moderate, marked, and severe. We then used the Wan2.2-Animate-14B animate video generation AI model to re-animate these modified portraits onto the original driving footage. This factorial design varied several model parameters including; pose retargeting, classifier-free guidance scales, and generation modes, resulting in 180 unique videos. Quality was measured through Frechet Video Distance (FVD) for distributional fidelity and a physics-aware assessment performed by a multimodal large language model to evaluate physical plausibility. Results: Our analysis yielded two primary observations. First, compositing through replacement-mode achieved significantly higher temporal fidelity than animation-mode (mean FVD 8.6 vs. 19.4; Cohen's d = 1.84). Second, while distributional fidelity showed a monotonic decline as hygiene perturbation increased (Spearman rho = 0.48, p < 0.001), physics-aware scores did not follow a similar trend. This pattern is consistent with fine-motor artifacts arising from model-level generative constraints rather than from the severity of the appearance modification alone. Conclusions: These findings demonstrate that generating appearance-modulated clinical video libraries is technically achievable. Nevertheless, the persistence of fine-motor artifacts underscores the necessity of expert human oversight before these materials can be safely deployed in educational and translational settings. Keywords: Generative artificial intelligence; Standardized patients; Video diffusion models; Psychiatric simulation; Mental status examination; AI-generated video; Medical education; Digital psychiatry
- Automated EEG Classification to Track Levels of Consciousness
Precise prognostication in acute brain injury is limited by a lack of reliable biomarkers of consciousness available to clinicians at the bedside. The ABCD framework is a method of classifying resting-state clinical EEG into categories that reflect levels of thalamocortical network function. ABCD classifications in the intensive care unit (ICU) have been shown to provide diagnostic and prognostic utility for patients with severe brain injuries, but the current gold standard for ABCD classification is visual inspection of power spectra, which is labor-intensive and requires expertise in spectral analysis. Using 4,611 manually classified EEG power spectra, we developed an automated, highly accurate, and well-calibrated convolutional neural net-based classifier of EEG into ABCD categories. The classifier has performance comparable to that of the current gold standard and that outperforms an alternative method of automated spectral analysis. As proof-of-principle for clinical implementation, we apply the classifier to a continuous EEG record from a patient with acute severe traumatic brain injury in the ICU, demonstrating its ability to yield continuous ABCD classifications that capture state fluctuations with high temporal and spatial resolution. The automated ABCD classifier allows for efficient analysis of continuous EEG records, facilitating the translation of the ABCD framework to the bedside for patients with acute severe brain injuries. The ABCD classifier also creates new opportunities to efficiently analyze large EEG datasets and generate new insights into the electrophysiological properties of human consciousness.
- Vobbin
Health & Fitness Coach
- Autoregressive Boltzmann Generators
Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has driven the development of Boltzmann Generators (BGs), which allow rapid generation of uncorrelated equilibrium samples by combining a generative model with exact li...
- Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing complex tasks into executable actions. While small open source MLLMs are cost efficient and privacy preserving compared with commercial large models, they suffer from...
- When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain is capped by a quantity the field rarely reports. For any policy whose output is one member model answer, accuracy cannot exceed one minus beta, wh...
- Dub Ninja
Live autonomous AI DJ that digs, mixes & explains 24/7
- Prompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection Settings
Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candidates to strategically manipulate algorithmic hiring systems. We study prompt injection in automated résumé screening, defined as subtle self-promotional text that introduces no new qua...
- Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decompose...
- Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text
Deep learning models have achieved impressive performance across various fields but remain vulnerable to adversarial inputs, particularly in NLP, where such attacks can have significant real-world consequences. Adversarial attacks often involve small, semantically similar token replacements to fool ...
- Automating Potential-based Reward Shaping with Vision Language Model Guidance
Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to correctly attribute the sparse success rewards to relevant parts of the trajectory. Naive reward shaping can induce reward hacking, yielding policies that exploi...
- TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the large number of visual tokens, which introduces substantial computational overhead. Visual token pruning offers a natural solution, yet existing methods are imperfe...
- OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed sympto...
- Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predicting discretized visual tokens. These tokens are derived from a codebook that maps embeddings to quantized visual patterns. The language-like architec...
- Joint Learning of Experiential Rules and Policies for Large Language Model Agents
For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or using trajectories ...
- Heavy-Ball Q-Learning with Residual Weighting Correction
This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes its convergence. It also identifies conditions under which the method is theoretically guaranteed to converge faster than standard Q-learning. The same construction is then extended to Q-lear...
- Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions
We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions. Building on the PINPOINT project and its use case, the EU Monitoring Mission in Georgia, we combine an interdisciplinary risk-model with OSINT-based media colle...
- Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning
Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone ...
- Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation
LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. We show that this can miss vulnerabilities introduced by fine-tuning itself: models can learn token-level indicator semantics that preserve canonical accuracy whi...
- NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem solving often requires not only factual knowledge but also quantitative reasonin...
- State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading
Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints. This makes the state representation given to a reinforcement learning agent an important design choice. We study this in HydroDam, a pumped-storage arbitrage...