AI News Archive: June 3, 2026 — Part 20
Sourced from 500+ daily AI sources, scored by relevance.
- Agentic Authoring of OMOP Concept Sets from Natural Language
Authoring OMOP concept sets from free-text descriptions remains a major bottleneck in scalable computable phenotyping for observational research. Existing tools support parts of this workflow but are designed primarily for interactive expert use rather than autonomous large language model (LLM) agents. We present an agentic framework that automatically generates OMOP concept sets by combining vocabulary tools, ontology extensions (RxClass, LOINC, and Disease Ontology), and procedural guidance. In ablation studies, the best configuration achieved Recall@100 of 0.965 and AP@100 of 0.875 on the development set. Cohort-level validation against OMOP-mapped EHR data yielded precision of 0.970, recall of 0.998, and a Jaccard index of 0.968. On an independent silver-standard benchmark of 457 concept-vocabulary pairs from 15 AD/ADRD target trial emulation studies, Recall@100 reached 0.835 and AP@100 reached 0.786. Task-specific tools outperformed unrestricted SQL access and PHOEBE 2.0, while progressive guidance performed best.
- Medication-Wide Association Study of Alzheimer's Disease and Related Dementias: Identifying Drug Candidates from Electronic Health Records through Explainable AI
Objective: Alzheimer's disease (AD) is a leading cause of death and disability, and treatment options for Alzheimer's disease and related dementias (ADRD) remain limited. We applied a data-driven, mechanism-agnostic Medication-Wide Association Study Plus (MWAS+) framework to identify candidate medications associated with ADRD using longitudinal electronic health record data and explainable artificial intelligence (AI). Methods: We used Veterans Health Administration electronic health record data from January 1999 to May 2022. The initial study population comprised 8,424,715 Veterans aged 65 years or older. Cases were defined by ADRD-related diagnosis codes or ADRD-related medication prescriptions, and controls were free of ADRD diagnosis and ADRD-related medication use. After exclusions and matching on sex, race, age at first encounter, and duration of follow-up, the primary analytic cohort included 505,817 matched case-control pairs (1:1; 1,011,634 Veterans). Longitudinal features were extracted from historical data up to 1 year before the index date and aggregated into 1-year intervals. We developed an upgraded Hybrid Value-Aware Transformer (HVAT 2.0) to jointly learn from longitudinal and nonlongitudinal clinical data while incorporating numerical values associated with clinical concepts, including cumulative medication dose. To enhance interpretability, we applied a medication-specific impact score method to estimate model-derived associations between medication exposure and ADRD risk. Findings: The model demonstrated stable performance across data partitions, with area under the receiver operating characteristic curve values of 0.791 in the training set, 0.772 in the validation set, and 0.775 in the testing set. Metolazone and varenicline were identified as the top 2 candidate medications with negative impact scores, suggesting potentially protective associations with new-onset ADRD. The impact score was -0.196 per unit of cumulative dose for metolazone (1800 mg) and -0.134 per unit for varenicline (280 mg). Although individual-level impact scores varied, most exposed patients had negative scores, including 12,020 of 12,480 metolazone users (96%) and 8,341 of 8,786 varenicline users (95%). Implications: This study demonstrates the feasibility of combining a medication-wide association framework, longitudinal dose-aware modeling, and explainable AI to identify candidate medications for ADRD from real-world electronic health record data. The findings should be interpreted as signals for hypothesis generation rather than evidence of causality. This framework may support prioritization of repurposing candidates for expert review, follow-up cohort validation, and future clinical investigation.
- Prognostic performance of an AI-based recurrence risk model in clinically low-risk HR+/HER2- early breast cancer
Objective Accurate prognostication of recurrence risk in HR+/HER2- early breast cancer is central for therapeutic decision-making, including identifying patients who may safely avoid adjuvant systemic therapy. However, the performance of existing prognostic tools remains insufficient for effective clinical stratification, motivating the development of artificial intelligence (AI)-based methods to improve risk stratification. Methods Ataraxis Breast CTX (ATX) is a multi-modal AI test that integrates H&E-stained whole-slide images with clinicopathologic features to predict risk of recurrence for individual patients. This study aims to validate ATX in an external dataset enriched for clinically low-risk patients from Dordrecht, the Netherlands. ATX scores were generated for 892 women diagnosed with early HR+/HER2- breast cancer. Of the 892 patients, 299 did not receive adjuvant systemic therapy. The discriminative performance of ATX was assessed using C-index and its stratification ability was evaluated by log-rank tests comparing Kaplan-Meier survival curves across risk groups. Results ATX achieved a C-index of 0.71 and a 5-year time-dependent AUC of 0.71, demonstrating strong discrimination in predicting recurrence-free survival (RFS). Among 299 patients who received no adjuvant therapy, ATX achieved a C-index and time-dependent AUC of 0.78 and 0.81 respectively, suggesting ATX retains prognostic information in the absence of systemic therapy. ATX scores were used to stratify patients into risk groups using a pre-specified threshold, where 656 (74%) were classified as ATX low-risk and 236 (26%) were classified as high-risk. Notably, untreated and treated ATX low-risk patients had comparable 5-year RFS (untreated: 5-year RFS = 96%, 95% CI = 92-97%; treated: 5-year RFS = 96%, 95% CI = 93-97%) with near identical 10-year RFS (86%, 95% CI = 83-92% for both), suggesting ATX low-risk status may identify a subgroup with favorable prognosis independent of treatment exposure. Conclusion ATX provides robust prognostic stratification in an external cohort of clinically low-risk HR+/HER2- early breast cancer and identifies a subgroup of patients who did not receive systemic therapy with favorable observed outcomes. These results support prospective validation of ATX as a decision-support tool for adjuvant therapy de-escalation in HR+/HER2- early breast cancer.
- Agentic Chart Review from Longitudinal Clinical Notes: a Lung Cancer Guideline Concordance Use Case
Clinical chart abstraction extracts structured patient variables from longitudinal clinical notes but is labor-intensive and difficult to scale. We evaluated LLM agents for question-guided chart review using lung cancer molecular testing guideline concordance as a use case. Two configurations were compared: (1) sequential note review using metadata and chronology, and (2) the same framework augmented with keyword-based note search. Gold-standard labels were established by human annotators. The search-enabled agent achieved higher accuracy (92.4% vs. 83.5%) and reduced errors by more than half (41 vs. 89) by retrieving evidence from long, heterogeneous note histories. In guideline concordance evaluation, most determinate patient-rule assessments were concordant (80.7%), while most apparent non-concordance reflected missing molecular testing documentation rather than documented care deviations. These results suggest tool-augmented LLM agents can approximate key aspects of human chart review and support scalable information extraction from longitudinal clinical documentation.
- Signal Quality Screening and Automated Sleep Stage Agreement in Home EEG: A Systematic Comparison of Dreamento and YASA on the Wearanize+ Dataset
Wearable EEG devices such as the Zmax headband offer scalable alternatives to laboratory polysomnography (PSG) for sleep monitoring, but their real-world performance in home settings remains poorly characterised. This study presents a systematic validation of automated sleep staging on the Wearanize+ dataset; a unique multimodal resource providing synchronised full PSG, bilateral Zmax EEG (F7-Fpz/F8-Fpz), and psychiatric phenotyping from 100 participants recorded at home. We first developed and applied an automated signal quality screening framework, revealing that 10% of recordings failed completely due to signal dropout and a further 16% showed partial degradation. We then evaluated two automated staging algorithms; Dreamento and YASA against PSG manual scoring, stratified by signal quality. In technically adequate recordings (N=74), YASA achieved significantly higher agreement than Dreamento (mean {kappa}=0.450 vs 0.371; {Delta}{kappa}=+0.079, p=0.0005), primarily through substantially improved N2 detection (recall: 0.64 vs 0.36). Both algorithms showed a systematic N2/N3 boundary confusion, however in opposite directions: Dreamento over-called N3 (37% of N2 epochs mis-staged as N3), while YASA over-called N2 (35% of N3 epochs mis-staged as N2). Critically, Dreamento showed greater robustness than YASA in degraded-quality recordings (WARN group: {kappa}=0.414 vs 0.330), consistent with its training on Zmax-specific data. Signal quality metrics did not predict staging performance within adequate recordings, indicating that channel topology is the primary limiting factor for frontal single-channel staging. These findings establish the Wearanize+ dataset as a benchmark for wearable sleep staging and motivate the use of PSG manual stage labels for downstream physiological analyses.
- Trump’s EO Furthers Model Exclusivity, Harming Cyber Defenders
The EO could strengthen relations between model providers and Washington, D.C., but comes with some gaps.
- Trump’s AI order gives Washington a look at frontier models, but not much leverage
The most powerful AI models are now treated, at least in Washington, as potential national-security events. Before companies release them to the public, the government wants a chance to see what they can do: whether they can discover software vulnerabilities, assist cyberattacks, or otherwise introduce risks that federal officials may not fully understand until the models are already in use. President Trump’s new executive order, signed Tuesday , is meant to give the government that chance. But the final version leaves AI companies with considerable control over the process. It asks them to voluntarily submit advanced models for government review 30 days before public release, and it does not make release conditional on what agencies find. That is a softer framework than the White House had been considering just last month. A previous draft had mandated a 90-day window, which tech industry executives opposed. The president nearly signed the first version of the order, but after a phone call with former AI and crypto czar David Sacks, the EO was put on hold. During another White House meeting on Monday, Sacks again stressed that longer wait times would stifle domestic development of AI models. The approach drew predictable praise from free-market groups. “The administration deserves credit for recognizing that innovation, not precautionary regulation, is what made America the global leader in AI,” says Competitive Enterprise Institute fellow Wayne Crews. The EO is careful to note that the government assessment program is voluntary for AI companies, and that public release of new models is not conditional on the outcome of the assessments. Given the potential destructive power of new AI models such as Anthropic’s Mythos, the order puts the government in a limited role: close enough to review the systems, but not necessarily empowered to slow them down, some tech policy analysts observed. Critics said the voluntary structure leaves too much power in the hands of the companies being reviewed. The consumer rights advocacy group Public Citizen called the arrangement a form of industry self-regulation, while the pro-regulation nonprofit Future of Life Institute argued that highly capable models such as Mythos require more than a “trust the companies” approach. “My impression is that it does not really establish the strong leadership that the federal government has traditionally had in terms of facilitating public-private partnerships and safeguarding responsibilities that have traditionally been left to the government like critical infrastructure,” Jessica Ji, senior research analyst at Georgetown’s Center for Security and Emerging Technology, tells Fast Company . The order does not prescribe a detailed testing regime. Instead, it sets up a framework and directs agencies to build the process. It calls on the National Security Agency and other security-focused agencies to co-design the model assessment framework and determine cyber-risk thresholds, especially around advanced cyber capabilities and what qualifies as a frontier model for the review regime. The Treasury Department will establish an AI cybersecurity clearinghouse to track the discovery and patching of software vulnerabilities exposed by new AI systems. Government agencies will use the 30 days for “cyber capability evaluations, adversarial testing, and national-security review” of large AI models, the EO states. The Commerce Department’s National Institute of Standards and Technology will play a key role, as will the Center for AI Standards and Innovation, formerly the AI Safety Institute, which already evaluates frontier models. Ji believes the influence of AI companies won’t end with the EO. “I’m personally very interested to see what this dynamic might look like in the future when it comes to who will lead on cybersecurity,” Ji says. “Do the AI companies get to set the terms as they release models, especially with this kind of weakened 30-day voluntary commitment to give the government access ahead of time?” In practice, many AI companies have already begun creating their own versions of early access and pre-release testing. Anthropic gave access to its Mythos model to a modest group of software and cybersecurity partners, and on Tuesday extended access to 150 new partners in more than 15 countries. OpenAI gave early access to its latest GPT-5.5 model to almost 200 trusted partners under its own early testing program, and a cybersecurity-focused version of the model remains available only to trusted partners. Those company-led programs may give some outside experts a look at the most capable new systems before they are widely released. But they also underscore one of the central tensions raised by the EO: whether the government can build an independent assessment process when the companies control much of the access, infrastructure, and technical information needed to evaluate the models. It’s also unclear whether 30 days is enough time for the government to properly assess the risks of an advanced AI model. “It depends on capacity to do evaluations, and I think the organizations best positioned to do those evaluations are the companies themselves,” Ji says. “So obviously we have a bit of a transparency problem: There’s this huge information asymmetry between the companies and everybody else, including the government.” The government might also face challenges in finding the right AI research talent and compute resources, as well as in managing access to the models and working out the details of the partnership with AI companies, Ji says. “I think a month probably does not mean that testers will have 30 days hands-on with the model,” she says. “It might look more like two weeks after they work through all the paperwork. It’s hard to say whether 30 days is adequate.”
- Knack
Agent Skills Management
- Vexavibes
Instant AI polls & surveys secured by a verifiable ledger
- Online Receipt Maker by FDM AI
Generate professional payment receipts instantly.
- QUATTRO — Four tiny tools, one window
AI answers, tasks and calendar without leaving your flow
- React App
Cryptographic identity and post-quantum crypto for AI agents
- AdSights Ads Framework
Programmatic ad production with Claude Code + Remotion
- TripAI
AI travel itineraries with interactive live maps in 30s
- Tarotool
Free AI tarot readings with clear 3-card insights
- Singify
Turn Any Text Into a Song with AI for FREE
- Palette Inspiration
Generate palettes inspired by 3,000 master painters
- Taxorio
Czech invoicing and VAT compliance made simple with AI
- Datailor Preference MCP
Give every AI agent your defaults
- Fiktion
AI Writing Workspace for Fiction Authors
- Pitch N Hire 2.0
The AI-Native ATS, Recruitment & Hiring Infrastructure
- VidGenn · Captions
AI-powered animated captions for videos in 38 styles
- Flowtrace
Watch and steer your AI agent as a live graph
- Rocketship
The only AI app builder with a built-in AI sales team
- FanQuiz
AI personality quizzes for creators, in any language.
- SpeakLearn
Speak English fluently with your personal AI tutor,Lea
- HowFast
The exact steps to do anything, powered by AI
- inbrowser.chat
Private AI chat that runs fully on-device in Chrome
- Strova
AI Standup Manager for Engineering Teams.
- Ones
One supplement, designed by AI from your blood & wearables
- Gen Pen by Obello
Refine designs and images faster with advanced AI editing
- Tubeup
Automate YouTube uploads — unlimited videos, AI Powered.
- UnsuitAi
Your AI-Powered Legal Assistant.
- Cooked
Track Claude Code context usage in real time with NPC roasts
- Basal
Know what your AI-built project costs before you build it
- Crukx
Ship reliable AI apps with autonomous testing agents
- Google will allow websites to opt out of AI overviews
Google will allow websites to opt out of AI overviews as pressure mounts over AI-driven traffic declines.
- Google must let British publishers opt out of AI search under new rules
Google must let British publishers opt out of AI search under new rules The Straits Times
- Trump signs order designed to give government early look at powerful AI models
Trump signs order designed to give government early look at powerful AI models The Washington Post
- Anthropic files for US stock market debut after valuation surge
Anthropic files for US stock market debut after valuation surge Computing UK
- Trump demands AI previews
Trump demands AI previews, Microsoft launches Scout personal assistant
- Introducing Gemma 4 12B: a unified, encoder-free multimodal model
Gemma 4 12B Unified Transformer
- High-order brain interactions distinguish wakefulness, anaesthesia, and recovery induced by deep brain stimulation
Understanding how consciousness depends on large-scale brain interactions is key for both the neuroscience of consciousness and clinical translation. However, it requires moving beyond classical pairwise descriptions of functional connectivity, which cannot capture the collective dependencies emerging across multiple brain regions. Here, we use multivariate information theory measures to characterize how higher-order interactions reorganize across states of consciousness. Specifically, we apply O-information to resting-state fMRI data from non-human primates to quantify whether multiregional brain dynamics are dominated by synergistic or redundant information sharing. We analyse two complementary datasets: (i) wakefulness and anaesthesia-induced loss of consciousness using different molecular agents (propofol, sevoflurane, ketamine), and (ii) the recovery of consciousness driven by central thalamic deep brain stimulation during propofol anaesthesia, indexed by behavioural responsiveness. We identify optimal regional subsets whose O-information robustly discriminates conscious from non-responsive states under two complementary optimization polarities. The first captures elevated redundancy in conscious scans that decreases under anaesthesia, providing robust discrimination and placing high-voltage central-thalamus stimulation closer to wakefulness. The second captures a synergy-to-redundancy transition, prominent in multi-anaesthesia conditions but context-dependent across datasets. Discrimination performance depends on interaction order: redundancy-based signatures improve with increasing subset size, whilst synergy-based signatures peak at low orders. Higher-order informational features significantly outperform pairwise functional connectivity, particularly for synergistic signatures which remain invisible to correlations. These findings demonstrate that consciousness is reflected in the reconfiguration of higher-order interaction structures, with distinct informational substrates requiring multivariate characterization beyond pairwise connectivity.
- Distance Mapping and Variable-Specific Geometry of Goal-Relevant Frames in the Retrosplenial Cortex
Goal-directed navigation requires animals to continuously update their position relative to an unmarked goal. Here, we recorded retrosplenial cortex (RSC) activity in freely moving rats during goal-directed navigation and random foraging. We found that RSC neurons encoded the Euclidean distance to the goal, and that this distance representation was selectively biased toward the goal during navigation. This goal-biased signal could not be explained by non-uniform behavioral sampling alone. Task engagement selectively enhanced allocentric head-direction representations anchored to a landmark cue, whereas egocentric boundary-bearing signals showed no detectable task-related enhancement and no detectable goal-centered spatial organization in this task context. Mixed-selective RSC population activity further exhibited variable-specific separability--smoothness geometry: distance-to-goal showed high local smoothness and decoding performance, whereas egocentric boundary bearing showed stronger macro-scale separability. These task-related spatial representations persisted under reduced visual input, suggesting contributions from memory and self-motion signals. Together, these findings indicate that RSC organizes goal-relevant spatial representations in a task-dependent manner.