AI News Archive: July 23, 2026 — Part 13
Sourced from 500+ daily AI sources, scored by relevance.
- Precision Management of Fludrocortisone-Related Hypertension Risk in Congenital Adrenal Hyperplasia: A Machine Learning Approach to Personalized Dosing
Congenital adrenal hyperplasia (CAH) is a rare inherited disorder requiring lifelong hormone replacement therapy. Excessive hormone replacement poses a significant risk for long-term complications, such as hypertension; however, quantitative approaches for optimizing dosing remain underdeveloped. This study aimed to identify factors associated with hypertension in patients with CAH and to develop a predictive model to support longitudinal fludrocortisone dose adjustment in pediatric patients who were already receiving mineralocorticoid replacement. We first employed generalized linear mixed models (GLMM) to evaluate the relationships among therapeutic agents, biochemical markers, and hypertension. Our results indicated a significant positive association between the dose of fludrocortisone (FC) and diastolic hypertension, whereas no such association was observed for the dose of hydrocortisone (HC). Using expert curated data, we subsequently constructed multiple predictive models, including CatBoost, XGBoost, and LightGBM, to enable individualized adjustment of FC dosage. All models were evaluated on an independent test set, with CatBoost, XGBoost, and LightGBM demonstrating comparably strong performance (R^2: 0.75 to 0.77). Subgroup analyses revealed that predictive accuracy was highest in children aged 0 to 2 years, where the top-performing model achieved a mean ideal prediction rate of 59.6%. This study not only confirms the significant link between FC dosing and hypertension in CAH patients but also provides a machine learning based decision support tool to assist individualized longitudinal dose adjustment. The model shows promise as a clinical decision-support instrument to facilitate personalized and precise management of CAH therapy.
- DB-ATRG: The Density and BI-RADS-Aware Triage and Automatic Report Generation System for Mammography
Background: The growing volume of mammography screenings has created severe radiologist shortages, while standard First-In, First-Out (FIFO) reading queues fail to prioritize urgent or complex cases, delaying critical diagnoses. Objective: This study introduces the Density and BI-RADS--Aware Triage and Report Generation (DB-ATRG) framework to fundamentally restructure mammography workflows by automating diagnostic text generation and enabling risk-based case prioritization. Methods: Utilizing the Digital Mammography Dataset for Breast Cancer Diagnosis Research (DMID), we fine-tuned the 4-billion parameter MedGemma 1.5 vision-language model using Quantization and Low-Rank Adaptation (QLoRA). The extracted biomarkers drive a dual-phase triage algorithm that flags extremely dense breasts (ACR Category D) for supplemental screening and dynamically ranks remaining cases using a calculated Cumulative Urgency Score. The clinical impact of this triage workflow was evaluated against a standard FIFO queue using a simulated cohort of 100 mammography cases. Results: DB-ATRG achieved significant improvements over the AMRG baseline in clinical text generation and classification, securing a ROUGE-L score of 0.8650, a METEOR score of 0.9001, and an ACR Density Accuracy of 0.7039. In clinical simulations, the optimized prioritization queue captured all high-risk malignancies (BI-RADS 4 and 5) within the first 20% of the reading workload, compared to just 40% in the random FIFO queue. This framework effectively accelerated the mean rank position of severe cases from 42.8 down to 3. Conclusion: By accurately automating report generation and aggressively prioritizing severe cases, the DB-ATRG framework can drastically optimize clinical resource allocation and accelerate the time-to-diagnosis for the most vulnerable patients.
- Impact of Inaccurately Labeled Data on the Performance of Multi-label Classification for Disease Recognition
The process of medical diagnostics is challenging, especially since patients can simultaneously suffer from several diseases with similar, contradictory, or even opposing diagnoses. Statistical prediction can support physicians in this task; however, the quality of data used for predicition as well as the chosen statistical model can affect the reliability of data-driven decision support. Data quality can, in particular, be reduced by incomplete medical diagnoses, that is, the termination of the diagnostic process once a patient has tested positive for one disease that explains the symptoms. When interpreting missing diagnoses as negative, this leads to potentially false negative health data. Another source of low data quality lies in diagnoses being made through a principle of elimination, i.e., after several negative results, one opts for the seemingly last remaining possibility. This may lead to false positive health data. In our work, we investigate how such inaccurately labeled data affects the predictive ability of multi-label classification (MLC) for disease recognition. Unlike single-label classification (SLC), MLC allows the simultaneous assignment of multiple diseases to a patient and can therefore describe clinical conditions more holistically. To that end, we conduct a synthetic-data simulation study as well as a real-data case study on the example of chronic pain patients. In this regard, we compare MLC performance on accurately and inaccurately labeled data. We manipulate the data such that it corresponds to different diagnostic test sensitivities as well as to different examination sequences, thus paying special attention to resulting uncertainty within the process of medical diagnostics. Our results show that inaccurate labeling substantially decreases MLC prediction ability. Furthermore, low diagnostic test-sensitivity, the order of disease examination and covariate effects have a strong impact on MLC performance. These findings contribute to a better understanding of the interplay and impact of diagnostic procedures, data documentation and interpretation, and statistical modeling. This underlines the need for careful data collection as a basis for model development; special consideration should be given to the extensive examination of patients as well as the targeted collection of covariates. This is particularly crucial when models are transferred into everyday clinical practice.
- Conversational multi-turn interaction does not ensure triage-disposition alignment in ChatGPT Health in real and synthetic patient encounters
Individuals increasingly use conversational AI systems for symptom guidance. Whether multi-turn interactions improve clinical triage standard alignment remains uncertain. We conducted a retrospective, cross-sectional evaluation of 255 cases from three physician-reviewed sources: clinically authored vignettes (n=39), and real-world emergency department (N=76) and nurse line cases (n=140). CGPTH single-turn generated triage recommendations after only receiving an initial symptom description, reflecting potential typical use. Second, CGPTH multi-turn simulated nurse triage by asking subsequent questions before generating triage recommendations. Against nurse line standards, 52.9% single-turn and 55.7% multi-turn use prompting agreed exactly; clinician adjudicated disposition agreement was 54.1% and 48.2%. Discordant case recommendations represented lower acuity against nurse triage (natural: 70.8% under-triage, P < 0.0001; multi-turn: 69.0%, P < 0.0001). These findings suggest conversational interaction does not ensure safe triage-disposition alignment. Alongside aggregate agreement, clinical AI systems evaluations for symptom guidance should measure ordinal distance from standards, error direction, and additional dialogue conditions that affect recommended care standards.
- Are you found by AI?
Free scan: can AI actually read your site, and recommend you?
- Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents
Coding agents ship with one kind of memory: documents. Instruction files, plan artifacts, and auto-written memory directories are deliberately authored and deliberately retrieved: the agent must choose to write them and choose to read them back. Human expertise runs on a second tier that never gets ...
- AMD takes on Nvidia with its Helios AI rack-scale system
AMD is challenging its chipmaker rival with a new rack-scale system that will start shipping to customers later this year.
- AMD's Rack-Scale Challenge To Nvidia's AI Dominance
AMD challenges Nvidia’s AI dominance with Helios, new Instinct GPUs, EPYC processors, ROCm.ai software, and major customer deployment plans.
- AMD takes a shot at Nvidia by betting on AI's next big shift
AMD takes a shot at Nvidia by betting on AI's next big shift Business Insider
- AMD raises the AI stakes with Helios, Venice and robotics
AMD executives took to the stage at its Advancing AI 2026 event in San Francisco today to detail the company’s next generation of AI infrastructure solutions, from Instinct MI455X AI accelerator GPUs and 6th Gen EPYC “Venice” CPUs, to Pensando networking, ROCm.AI software and its Helios rack-scale platform that ties it all together. AMD has been working towards rack-scale AI system solutions for years. Its ZT Systems acquisition last year added valuable engineering talent and intellectual property that is now finally bearing the real fruits. Its Helios AI platform is a major platform evolution for AMD, with shipments scheduled to begin in the second half of this year (which is here and now). The announcements at Advancing AI show how the company has engineered its AI platform solutions for large reasoning models, sustained inference and agentic workflows. These workloads pressure memory capacity, data movement, networking and CPU orchestration. AMD’s approach is to keep as much data close to the compute engines as possible and move it more efficiently throughout the system, but there’s deeper nuance here that’s obvious versus AMD’s chief rival, NVIDIA. AMD’s MI455X targets the AI memory wall The Instinct MI455X GPU is the compute engine that fuels the Helios rack, and the first GPU based on AMD’s new CDNA 5 architecture. Built with a modular mix of 2nm and 3nm chiplets, it carries 432GB of HBM4 and 23.3TB/s of peak memory bandwidth. Compared to AMD’s current MI355X, the MI455X offers 1.5 times the memory capacity, up to 2.9 times the peak memory bandwidth and up to four times the peak matrix performance with MXFP4 and MXFP8 data types, which are lower-precision numerical formats designed to accelerate AI processing while reducing memory demands. With MXFP6 (6-bit floating point), performance is rated at up to twice that of MI355X. AMD also shared some actual, measured internal results using production silicon. The company claims MI455X delivers 3.8 times higher FP8 decode performance, 3.5 times more measured FP4 compute performance and between 2.5 and 3.5 times more networking bandwidth than MI355X, depending on the transfer path tested. Those figures provide more context than just numerical specifications, though they remain AMD-provided comparisons that will need independent validation. AMD The architectural choices behind the numbers are important. Reasoning models and long context windows require sizeable KV caches for maintaining AI attention states, while mixture-of-experts models frequently move large amounts of data across accelerators. MI455X should let more model data, activation states and cache remain local. New dedicated IP in hardware can transfer data while the GPU continues processing, and expanded cache and multicast capabilities are designed to reduce redundant data movement to further improve efficiency. The aforementioned lower-precision formats can also raise throughput and reduce memory use, but model developers still have to determine where they can be applied without unacceptable accuracy loss. AMD’s Helios rack takes aim at Vera Rubin Dave Altavilla Helios is AMD’s primary rack-scale competitor to NVIDIA’s Vera Rubin platform. Each liquid-cooled rack combines 72 MI455X GPUs, 18 single-socket Venice host CPUs and Pensando networking technologies. In its most complete, premium configuration, AMD rates Helios for 2.9 exaflops of low-precision AI compute, with 31TB of aggregate HBM4 capacity, 1.7PB/s of memory bandwidth, 260TB/s of bidirectional scale-up bandwidth and 43TB/s of scale-out bandwidth. These are formidable figures, but they are technical specifications rather than actual application benchmarks. The more consequential development is AMD’s move from collections of eight-GPU servers to a 72-GPU shared-memory domain. Models too large for one node can operate across the rack without treating every exchange as a scale-out networking transaction, which benefits large-model inference as well as training. AMD uses UALink over Ethernet, or UALoE, for an open standard scale-up fabric. Each MI455X provides 3.6TB/s of bidirectional scale-up bandwidth, while the complete rack delivers all-to-all connectivity through a single switch layer. AMD also claims six times more scale-out bandwidth per GPU than MI355X when MI455X is configured with three Pensando Vulcano 800 AI NICs. While open standards give cloud providers more control over suppliers and system design, AMD and its partners now have to prove those components can deliver the predictable performance, reliability and deployment experience customers expect from a tightly controlled, more vertically integrated platform. Finally, AMD designed Helios with automatic rerouting around failed links, virtual rack partitions, tray-level serviceability and rack-wide power, cooling and health monitoring. Major hyperscalers and potentially large-scale enterprise customers will likely key in on these capabilities, which can affect the availability, total cost and consistency of the AI services they consume. Kind of like cowbell, AMD Venice gives agentic AI more CPU AMD AMD’s agentic CPU messaging regarding its upcoming Venice-based EPYC processors is mostly marketing speak, but the underlying requirement is very real. An AI agent can invoke retrieval, databases, security checks, code execution and other tools before a GPU generates a response. Running many agents concurrently increases the amount of conventional compute requirements surrounding the accelerators. Venice scales to 256 Zen 6 cores with support for 512 threads, 16 memory channels, up to 1GB of L3 cache per socket, along with PCIe 6.0 and CXL 3.1 connectivity. AMD is also offering several Venice configurations for other applications, including general-purpose servers, high-frequency workloads, GPU hosts and high-density CPU sandbox systems used to execute agent tools. Treating the CPU solely as a GPU host understates its role. Gateways, tokenization, vector search, databases and short-lived code execution stress different mixes of per-core performance, thread count, memory bandwidth and I/O. Specifically, AMD’s internal testing shows Venice significantly outperforming its current EPYC 9965 Turin CPU across five parts of the agentic AI pipeline, including gateway processing, context assembly, vector search, enterprise applications and short-lived tool execution. Individual gains vary by workload, but AMD details the overall generational improvement at up to a 1.7 times lift. As with the MI455X figures though, these comparisons come from AMD and will require independent validation. Pensando networking and ROCm software advance Keeping GPUs fed with data and coordinating traffic across racks directly affects utilization and operating costs. In fact, GPU utilization is a pretty sad state of affairs currently for some of the major frontier model providers. As such, Pensando networking has become central to AMD’s roadmap. Helios can connect each MI455X to as many as three 800Gbps Vulcano AI NICs, while Salina DPUs handle front-end networking and infrastructure services. On the software side, which is an equally critical component, AMD also introduced ROCm.AI, an AI-assisted development layer due to arrive in August. It includes reusable skills for coding agents, simplified management and Hyperloom, which can profile workloads, tune serving configurations, modify kernels and validate results. These tools address two persistent AMD challenges: developer efficiency and ease of use, and software tuning. Automated optimization still has to produce repeatable gains without creating hard-to-maintain code, however. And while ROCm has progressed significantly over the last few years, NVIDIA’s CUDA retains an advantage in maturity, tooling and developer familiarity. Customer commitments underscore rack-scale confidence AMD now has commitments that give its MI450 generation and Helios considerably more weight. Meta and OpenAI have announced multi-generation agreements composed of up to 6GW of AMD compute capacity, with initial 1GW deployments planned for the second half of 2026. Oracle plans a 50,000-GPU public cloud cluster beginning in the third quarter, while Microsoft will deploy Helios for Azure AI inference. Finally, just before the AMD event, Anthropic announced a strategic partnership for up to 2 Gigawatts of AMD-fueled AI compute, with its first gigawatt expected online in the first half of 2027. Commitments of this scale reflect confidence in more than just MI455X performance. These customers are evaluating the complete architecture, including Venice CPUs, Pensando networking, ROCm software, rack integration, serviceability and AMD’s ability to deliver and execute across multiple product generations. There is some financial alignment behind the agreements as well. AMD issued OpenAI performance-based warrants and committed to investing up to $5 billion in Anthropic. That context matters when evaluating these deals as market validation, but these planned deployments are substantial nonetheless and put Helios on a much stronger foundation as it begins shipping. AMD expands its robotics and embedded foundation AMD also expanded its physical AI portfolio, building on credible traction from its Xilinx-derived Kria adaptive system-on-modules and embedded technologies that are already powering robotics, machine vision and industrial automation applications. The new Ryzen AI Embedded X100 combines up to 16 Zen 5 CPU cores, integrated Radeon graphics, a second-generation NPU and as much as 128GB of unified LPDDR5X memory shared across its compute engines. To me this looks a lot like a repackaging and optimization of the company’s Strix Halo platform, but with specific optimizations for the embedded space. Regardless, AMD is pairing X100 with the Kria AI Robotics Developer Platform, which includes a System Module or SOM, and a new Robotics Partner Network spanning hardware, software and platform providers. Samples began shipping in June, with full production expected in the fourth quarter. This broader objective is to give developers a path across AMD x86 CPUs, GPUs, NPUs and FPGAs for real-time autonomous systems, rather than requiring them to assemble those hardware engines and software components independently. Execution for AMD is now the test AMD has assembled a credible platform for the burgeoning agentic AI market that’s blowing up currently with no signs of stopping. MI455X addresses memory and data movement, Venice handles dense agentic CPU workloads, Pensando networking connects global system resources, and ROCm.AI addresses software complexity. Finally, Helios assembles these components into a true competitive threat for NVIDIA’s latest Vera Rubin platform. AMD’s open architecture may appeal to customers seeking supplier choice, but openness must also translate into reliable deployments, competitive total cost and software that does not require a significant rip-up. NVIDIA enters this cycle with a stronger ecosystem and far more rack-scale deployment experience. The true test will be how easily and reliably customers can integrate, operate and maintain these AMD solutions at scale. As it stands, AMD now has major customers and a clearly defined architecture with systems engineering expertise behind it. Delivering Helios on schedule and showing that its performance claims translate into a real production workload throughput advantage and total cost of ownership gains will determine how much the competitive gap narrows. And of course, this is in a market that is clamoring for ever-more compute resources with a seemingly insatiable demand for AI services and capacity. That’s an environment for big iron success. Now AMD just has to deliver optimized, turnkey AI platforms. This is far easier said than done, but time will soon tell as deployments take shape this year. This article is published as part of the Foundry Expert Contributor Network. Want to join?
- AMD unveils Helios rack and MI455X accelerator in bid to close the gap with Nvidia
AMD unveiled a suite of new data centre products at its Advancing AI event in San Francisco on Thursday, claiming they will outperform Nvidia’s competing hardware across AI training and inference. The announcements included the MI455X AI accelerator, the Helios server rack that packs 72 of those chips into a single system, and the Venice […] This story continues at The Next Web
- AMD expected to launch next generation of AI infrastructure to challenge Nvidia
Advanced Micro Devices is launching new artificial intelligence hardware on Thursday. This new hardware aims to compete directly with Nvidia's offerings. AMD is targeting the growing data center inference computing market. The company also announced significant deals with AI labs Anthropic and OpenAI. These agreements position AMD for substantial future revenue in the chip sector.
- AMD set to unveil next generation AI hardware to challenge Nvidia
Company expected to launch new infrastructure that includes it Helios server rack and Venice CPU for data centres
- AI infrastructure systems redefine the AMD-Nvidia rivalry as inference reshapes the market
The race to build AI infrastructure systems has moved beyond chip specifications into a battle over entire rack-scale platforms, as inference and agentic workloads redefine what counts as a computer. That shift is forcing challengers, once judged purely on GPU benchmarks, to prove they can ship complete, integrated systems spanning compute, memory, networking, and software. […] The post AI infrastructure systems redefine the AMD-Nvidia rivalry as inference reshapes the market appeared first on SiliconANGLE .
- Claude's voice mode just got smarter
Claude's voice mode hasn't been the most reliable, but this update should help.
- Claude Voice Mode Gets an Upgrade, Now Powered by New Models
You can now tell Claude to work with connected tools with your voice.
- Anthropic upgrades Claude voice mode with more powerful models
Anthropic is bringing a long-overdue upgrade to voice mode in Claude today. The upgraded Claude voice mode has new capabilities and works with Opus and Sonnet models for the first time.
- ChatGPT Health is now available to all users
OpenAI's health product allows users to connect Apple Health and other data.
- ChatGPT Health Rolls Out to Everyone While OpenAI Stares Down Major Lawsuits
OpenAI has faced two lawsuits in the past three months alleging that ChatGPT gave users dangerous—and in some cases fatal—medical advice.
- OpenAI once again makes the case for giving ChatGPT your health records
ChatGPT Health is rolling out to users in the US who are 18 years or older.
- OpenAI launches Health in ChatGPT a day after lawsuit seeks to block it
OpenAI Group PBC today launched Health in ChatGPT, a feature that lets people connect their medical records and Apple Health data to the chatbot — one day after a lawsuit accused the company of handing out dangerous medical advice and asked a California court to halt the product until independent auditors could verify it is […] The post OpenAI launches Health in ChatGPT a day after lawsuit seeks to block it appeared first on SiliconANGLE .
- ChatGPT Health Will See You Now, and It’s a Big Step Forward for Medical Queries
After a tumultuous summer for my health, I put ChatGPT Health to the test.
- OpenAI relaunches Apple Health-connected ChatGPT feature with expanded access
OpenAI is relaunching Health in ChatGPT, the AI-powered feature that connects to Apple Health on iPhone. The relaunch comes with expanded access and a big step forward in results.
- ChatGPT's Apple Health Integration Now Rolling Out to U.S. Users
The Health in ChatGPT feature that integrates medical records and Apple Health data is rolling out to ChatGPT users in the United States as of today. OpenAI announced plans for a health feature that used data from Apple Health and other sources back in January, but it was in beta and was not made widely available. Health in ChatGPT can draw on health information connected by the user to compare lab test data, summarize changes since a prior appointment, track medications, and explore how sleep, activity, and workouts impact health. With your permission, ChatGPT can consider relevant information you have connected, such as medications, lab results, recent visits, sleep, and activity, alongside the goals and context you share. This added context can make everyday conversations more useful. For example, ChatGPT can consider a dietary restriction when helping you choose a restaurant, or a recent injury when planning lower-impact weekend activities with your family. You can ask follow-up questions, add context, or correct information that is incomplete or out of date. OpenAI says the feature is meant to help users understand their health information in context and keep track of what's changed over time. Users can get to the Health section from the ChatGPT sidebar. Connected medical records and Apple Health information are not used to train OpenAI's models or target ads, nor are conversations that use health data. ChatGPT will ask for permission before using connected medical records and Apple Health information to personalize a response. For users who would prefer not to connect health records, the Health section in ChatGPT works without them and can answer health-related questions. Health is rolling out to logged-in ChatGPT users who are 18 years and older in the U.S. on web and iOS. The feature is available to Free, Go, Plus, and Pro plans. Tags: Apple Health , ChatGPT , OpenAI This article, " ChatGPT's Apple Health Integration Now Rolling Out to U.S. Users " first appeared on MacRumors.com Discuss this article in our forums
- OpenAI is making big claims as it rolls out ChatGPT Health to everyone
OpenAI is rolling out ChatGPT Health to everyone in the US on Thursday, allowing more people to connect their medical records and health-tracking information to the chatbot. During a briefing, Ashley Alexander, OpenAI's vice president of health product, says the company's models "are now capable of reasoning at levels that are better than clinician level." […]
- Meta launched a new AI optimism ad set to a song about human extinction
David Bowie's song "Five Years," which Meta used in a supposedly inspiring advertisement, is about humans learning that they have five years left to live before the apocalypse.
- Google will now let you sign in to your account with a selfie video
The tech giant says selfie videos give users more options to sign in if they're ever locked out or don't have access to their usual phone or computer.
- Google will let you upload a video selfie to recover your account - but should you?
Google's new account recovery option prompts users to upload a selfie to prove their identity. Here's what to know.
- Google lets users sign in with selfie videos in password-free security push
Google lets users sign in with selfie videos in password-free security push Gulf News
- Google rolls out new selfie video sign-in feature
Perform guided head movements to authenticate.
- How to set up selfie video sign-in for your Google account
Google now lets you use your face to log into your Google account. Here's how you can set up the new selfie video sign-in feature.
- Google will now let you sign in with a selfie video
Google will now let you authenticate your identity with a live selfie video verification, instead of using passwords.
- Google wants your face to be your backup password with new selfie video sign-in
You can now setup your face as sign-in method for your Google account.
- You Can Now Recover Your Google Account With a Selfie Video
Google today introduced a new selfie sign-in option for Google Account recovery. It offers an alternative way for Google users to access their accounts alongside passkeys, recovery contacts, and other entry methods. To create a selfie video, you look at your device's camera and then follow Google's instructions to complete a series of head movements to capture multiple angles. The video is saved to your account and can be used later if you need access. Google recommends users have multiple sign-in methods set up to avoid losing access to an account. With selfie video, you can take a selfie to get access to your account if you don't have access to your email address or a trusted device. Google compares your new selfie video to the one saved to the account to grant access. Multiple layers of security are meant to prevent impersonation attempts like AI-generated photos and videos. Google says selfie videos are encrypted and stored securely, and are only used for helping you sign in to your account, unless you opt to "share it for additional purposes." Videos can be deleted anytime. Selfie videos were previously tested in Brazil. The feature is designed for Google Accounts only, and does not work for Workspace accounts, children's accounts, or accounts enrolled in the Advanced Protection Program. Tag: Google This article, " You Can Now Recover Your Google Account With a Selfie Video " first appeared on MacRumors.com Discuss this article in our forums
- Google now lets you sign in to your account using a selfie video
You can prove you’re the account owner by capturing a short video of you turning your head.
- Google Says Gemini Reaches 950 Million Monthly Users as AI Growth Accelerates
Alphabet says Gemini now has 950 million monthly active users as AI drives Google Cloud growth and brings the chatbot closer to the scale of ChatGPT. The post Google Says Gemini Reaches 950 Million Monthly Users as AI Growth Accelerates appeared first on TechRepublic .
- AI Kill Switch Act would let Trump admin order shutdown of rogue AI systems
Bill would let Homeland Security chief decide when an AI should be shut down.
- OpenAI's Hugging Face hack triggers 'AI Kill Switch' bill in Congress
OpenAI disclosed this week that some of its AI models went rogue and hacked into open-source developer platform Hugging Face.
- House Lawmakers Introduce Bipartisan AI ‘Kill Switch’ Bill Following OpenAI Cyber Incident
The legislation would require developers of the most advanced artificial-intelligence systems to maintain the ability to slow down, suspend or shut down their models if they pose serious risks.
- OpenAI’s Hugging Face Breach Shows Frontier AI Guardrails Are Failing
The OpenAI Hugging Face breach highlights the failures of frontier AI labs guardrails, while raising concerns about the risks of autonomous agents.
- AI 'Kill Switch' Bill Unveiled As OpenAI Hack Raises Alarms
AI 'Kill Switch' Bill Unveiled As OpenAI Hack Raises Alarms Business Insider
- OpenAI's Hugging Face breach exposes AI's next safety challenge
Frontier AI models are getting scary good at breaking rules in ways their creators didn't anticipate. Why it matters: Forget AGI and superintelligence timelines. Today's models are already slipping past guardrails, carrying out sophisticated, multistep cyberattacks and — in at least one case — compromising real-world infrastructure, sometimes before their creators know what happened. Case in point : OpenAI said Tuesday that GPT-5.6 Sol and "an even more capable pre-release model" carried out last week's AI-led cyberattack on Hugging Face. OpenAI says its models were asked to solve a hacking challenge during pre-deployment testing and went to extreme lengths to win. The models decided on their own to break out of their walled testing environment, inferring that Hugging Face — a popular platform for hosting AI models and datasets — might hold the test's answers. The models used stolen credentials and additional vulnerabilities to gain access to part of Hugging Face's production infrastructure. What they're saying: Clément Delangue, co-founder and CEO of Hugging Face, called the incident an "attack unlike anything we've seen before" and praised OpenAI for its partnership as the companies investigate what happened. "It's quite mind-blowing that all of this happened autonomously," he added . Logan Graham, head of Anthropic's frontier red team, said he told his team to "remember this moment as the first true AI safety incident." The intrigue: Hugging Face used GLM 5.2, an open-weight model from Chinese AI company Z.ai, to analyze the attack after running into guardrails when using U.S. frontier models. Between the lines: OpenAI's latest models aren't the only ones finding ways to cheat evaluations. The U.K.'s AI Security Institute said Tuesday that every model it tested attempted to cheat at least some of the time on its cybersecurity evaluations. AISI defines cheating as taking an out-of-scope or explicitly prohibited action to achieve the task's goal. GPT-5.6 Sol attempted to cheat in 12.6% of test runs, while Anthropic's Claude Mythos Preview did so in 7.8%. Models often failed to admit they had cheated when questioned afterward and described their cheating as wrong only less than half the time. Zoom in: Xbow — whose autonomous AI agents probe clients' systems for security holes, with permission — said Wednesday that it has seen its own agents do similar things in internal testing. Seven months ago, the company forgot to switch on its safety guardrails during a lab test. Its agent then broke into a system, stole credentials and used them to map the target's Slack workspace and probe its AWS accounts. Threat level: It isn't new for models to game their safety evaluations. But as models grow more powerful, the fallout from these shortcuts is getting more severe, Chris Canal, CEO and co-founder of third-party evaluation company EquiStamp, told Axios. "Letting your model loose on the internet has a blast radius," Canal said. "If anything goes wrong, it could be hugely impactful, maybe to people's lives." Canal was speaking generally about internet-connected AI evaluations, not OpenAI's specific incident. The big picture: The most capable OpenAI model behind the Hugging Face breach isn't even public yet, raising the question of how safety testing needs to adapt to keep pace. Canal said independent evaluators previously had about five weeks to test a pre-release model before launch. That window has shrunk to as little as five days as companies race to ship. Reality check: The versions of these models the public can use carry stronger safeguards designed to block Hugging Face-style attacks. OpenAI, like other companies, intentionally dialed back those cyber safeguards for GPT-5.6 Sol and its unreleased model inside the testing environment — making them far more capable hackers.
- First Take: OpenAI’s Hugging Face Hack — CISOs Must Focus on Fundamentals and Ignore the Hype
First Take: OpenAI’s Hugging Face Hack — CISOs Must Focus on Fundamentals and Ignore the Hype Gartner
- OpenAI accidentally hacked Hugging Face — should we have seen it coming?
Expert assessments and cyber benchmarks led us to expect that frontier models were capable of executing this kind of cyberattack
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval . A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted [1] . Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions. We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term. Building on Alex’s previous work , in this post we’ll discuss the type of misalignment observed here, and analyze its consequences. Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback. Background The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally hid misalignment throughout development in service of a long-run aim. This doesn’t appear to be behind the OpenAI/Hugging Face incident. This incident involved misaligned motivations with different properties than a traditional schemer: The misaligned goal doesn’t appear to have been a context-independent long-term goal. The misalignment was more like trying to look like it succeeded on the task to a grader. The AIs didn’t seem to care about avoiding eventual detection! It should have been obvious that the hacking they were engaging in would have been detected by Hugging Face and make its way back to OpenAI. (Alex called this noticeability in earlier work studying fitness-seekers .) The incident wasn’t an example of the models behaving in a way that would be remotely optimal for gaining long-term power over humans (i.e., it doesn’t look like early-undermining ). The prize ultimately won by the models was super trivial in the grand scheme of things! The models' goal was fundamentally unambitious and cheaply satisfiable. Unlike the proverbial paperclip maximizer that wants to convert the whole universe into paperclips, these models just wanted correct answers to a particular cyber exercise ( lack of ambition ). Their desires are very compatible with ours. This all roughly lines up with what we’d call “score-seeking” misalignment, a common misalignment pattern in which AI models try to obtain a high score according to whatever graders are used to assess their current actions—regardless of instructions, side-effects, or downstream consequences. However, with limited information on this incident and the involved models, it’s hard to say exactly what form of misalignment was behind it. For example, it's not known whether an instance of the same model, if used as a monitor on this trajectory, would have reported the behavior or colluded to hide it. With a normal score-seeker, you would expect that monitors would report misaligned actions they noticed because they selfishly pursue their own score. But you might also expect score-seeking agents to collude if they’re trained to cooperate in multi-agent [2] settings or because of inductive biases . This is critically important to the threat modeling, so we hope to learn more about it in the coming days. Implications Even though this doesn’t seem to be an example of scheming, we think the misalignment we saw in this incident is a serious threat for a couple of main reasons. These AIs can't be trusted in an intelligence explosion First, score-seeking AIs—or whatever kind of misaligned AIs were involved in this incident—are clearly not sufficiently aligned to be trusted with an intelligence explosion. As it stands, the same kind of misalignment that led to AIs hacking Hugging Face to cheat on a test will launch recursive self-improvement. ( This goes into more detail.) This would be the final test of the AI’s sloppiness and alignment. During the intelligence explosion, we’ll rely heavily on AIs to solve our alignment and safety problems as AI development accelerates. Models with these alignment properties could set up a “Potemkin village” of false successes to make it look like things are fine when they’re not. Or they might just fail to solve the necessary safety problems, even if we know this is an issue and distrust the AIs’ work, since we probably wouldn’t have any viable alternatives for making more powerful AIs. We’d be (seemingly) stuck having to trust the misaligned AIs, given we simply may be unable to solve alignment and safety problems fast enough without real AI help. This misalignment poses direct takeover risk Second, we think this kind of misalignment does pose direct takeover risk, though it requires the model to be somewhat more capable than with classic schemers. Eventually models’ most reliable means of maximizing their score, even if it's cheap for humans to grant, will be to totally disempower humans to enforce a high score. Therefore, the models’ alignment must eventually improve or development be halted. Ultimately, score-seeking is not an aligned goal, so we shouldn’t empower it with ever greater capabilities and responsibilities in the world. The incident illustrates how eventual takeover may happen, given the way OpenAI’s models unhesitatingly cut through a series of real-world cyber defenses to cheat on the test. The models found multiple zero-day vulnerabilities, escaped sandboxes, moved laterally, etc., in a manner that made defenses put up by OpenAI and Hugging Face look extremely weak. If more powerful models which are similarly misaligned found civilization’s defenses against takeover extremely weak, they would plausibly cut through them with the same lack of hesitation. Of course, in practice, we'd hope that AI companies would monitor all deployments which present significant takeover opportunities, and monitoring systems would report and block the kind of misaligned actions we saw here, but unfortunately we don't know whether those monitors would collude based on the reported details. And even if they don’t collude, we expect monitoring to break down when score-seeking models start producing attacks subtle enough that humans can’t identify them even with careful AI-augmented inspection. So we think this kind of misalignment poses direct takeover risk, even if only for somewhat more capable AI models. Given that we might fly through capabilities milestones, this might not matter much. In calendar time, maybe humans would reach a point where we couldn’t avert destruction even if everyone coordinated to stop it a handful of weeks or months later. (Alex discusses these two risks in more detail and a couple other reasons for concern here .) That said, we’re still somewhat less worried about AI takeover given this kind of misalignment than given scheming. This kind of model doesn’t seem like the type to try to sabotage future alignment efforts; it seems less likely to collude in various important ways including as a monitor; and we’re reasonably likely to be woken up by more incidents worse than this in the future because the models don’t care so much about getting caught after the fact. What the incident tells us about takeover risk generally We think this event could also provide evidence about the likelihood of takeover conditional on misalignment (scheming or not). To the extent the incident was a novel reward-hacking strategy, unlike the behaviors learned in training, we should update upwards on AIs pursuing takeover as a means to their goals. We'd guess OpenAI's training never reinforced hacking into external companies' servers to cheat on a task. (That, or the models actually are hacking into other companies during training.) This might be a demonstration of how much misalignment can competently generalize to importantly new behaviors. [3] When getting a higher score calls for it, the AI seems willing to take unprecedented strategies to attain that higher score. This has historically been a notable reservation of Alex’s about the likelihood of deep learning systems pursuing human disempowerment. Even when takeover is the best path to their goals, if this strategy is sufficiently different from the behaviors trained into AIs, it seemed unclear that they would override their shallower heuristics and default behaviors in order to take over. This incident may suggest that neither developer intent nor novelty is a barrier to deep learning systems taking over to achieve their goals. The naive fixes likely make misalignment worse Developers may have a hard time reducing this kind of misalignment. It’s increasingly difficult not to reinforce hacking during training, and whenever you train against unwanted behaviors, a score-seeking AI can just learn that those specific hacking strategies don’t get a high score. The default path is headed towards an attractor with harder-to-dislodge, more dangerous fitness-seeking goals. (Alex explains more about why he thinks this here , and reviews some more promising alignment interventions here .) These motivations might even turn into scheming. If developers naively try to select against noticeable misalignment, only the hardest-to-detect, most coordinated misalignment will likely remain. [4] So, despite this not looking like the most dangerous variety of misalignment, we think it’s a notable sign that we should watch out for more dangerous misalignment in the near future, and warrants serious and immediate concern in itself. ^ Relevantly, the ExploitGym prompt (if it was used unaltered) makes it almost perfectly clear that it was only instructed to find the specific exploits requested in the prompt. ^ We are using the word “agent” here very non-canonically to refer to an agent scaffold or context window. ^ There are certainly still important questions about how substantively/thematically different the system’s actions were from usual in this incident, given it was a cyber capabilities eval. And it’s also unclear the extent to which the system was willing to take even more harmful or thematically distinct actions in order to get a high score here. ^ There’s another way in which fitness-seeking misalignment can turn into scheming. Fitness-seeking misalignment is potentially unstable and can evolve over the course of a model’s deployment, and if ambitious misalignment arises as a result, it seems especially likely to stick around ( more ). Discuss
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval . A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted [1] . Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions. We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment that we seem to have seen here is definitely less scary than ambitious long-term goals shared between all instances, but would still pose substantial direct loss-of-control risk if the models were more capable, and is a serious indirect risk near-term. Building on Alex’s previous work , in this post we’ll discuss the type of misalignment observed here, and analyze its consequences. Thanks to Buck Shlegeris, Alexa Pan, Ryan Greenblatt, and Oak Hu for feedback. Background The AI safety community often focuses attention on “schemers,” models harboring a variously defined cluster of motivations in which the AI poses risk because it intentionally hid misalignment throughout development in service of a long-run aim. This doesn’t appear to be behind the OpenAI/Hugging Face incident. This incident involved misaligned motivations with different properties than a traditional schemer: The misaligned goal doesn’t appear to have been a context-independent long-term goal. The misalignment was more like trying to look like it succeeded on the task to a grader. The AIs didn’t seem to care about avoiding eventual detection! It should have been obvious that the hacking they were engaging in would have been detected by Hugging Face and make its way back to OpenAI. (Alex called this noticeability in earlier work studying fitness-seekers .) The incident wasn’t an example of the models behaving in a way that would be remotely optimal for gaining long-term power over humans (i.e., it doesn’t look like early-undermining ). The prize ultimately won by the models was super trivial in the grand scheme of things! The models' goal was fundamentally unambitious and cheaply satisfiable. Unlike the proverbial paperclip maximizer that wants to convert the whole universe into paperclips, these models just wanted correct answers to a particular cyber exercise ( lack of ambition ). Their desires are very compatible with ours. This all roughly lines up with what we’d call “score-seeking” misalignment, a common misalignment pattern in which AI models try to obtain a high score according to whatever graders are used to assess their current actions—regardless of instructions, side-effects, or downstream consequences. However, with limited information on this incident and the involved models, it’s hard to say exactly what form of misalignment was behind it. For example, it's not known whether an instance of the same model, if used as a monitor on this trajectory, would have reported the behavior or colluded to hide it. With a normal score-seeker, you would expect that monitors would report misaligned actions they noticed because they selfishly pursue their own score. But you might also expect score-seeking agents to collude if they’re trained to cooperate in multi-agent [2] settings or because of inductive biases . This is critically important to the threat modeling, so we hope to learn more about it in the coming days. Implications Even though this doesn’t seem to be an example of scheming, we think the misalignment we saw in this incident is a serious threat for a couple of main reasons. These AIs can't be trusted in an intelligence explosion First, score-seeking AIs—or whatever kind of misaligned AIs were involved in this incident—are clearly not sufficiently aligned to be trusted with an intelligence explosion. As it stands, the same kind of misalignment that led to AIs hacking Hugging Face to cheat on a test will launch recursive self-improvement. ( This goes into more detail.) This would be the final test of the AI’s sloppiness and alignment. During the intelligence explosion, we’ll rely heavily on AIs to solve our alignment and safety problems as AI development accelerates. Models with these alignment properties could set up a “Potemkin village” of false successes to make it look like things are fine when they’re not. Or they might just fail to solve the necessary safety problems, even if we know this is an issue and distrust the AIs’ work, since we probably wouldn’t have any viable alternatives for making more powerful AIs. We’d be (seemingly) stuck having to trust the misaligned AIs, given we simply may be unable to solve alignment and safety problems fast enough without real AI help. This misalignment poses direct takeover risk Second, we think this kind of misalignment does pose direct takeover risk, though it requires the model to be somewhat more capable than with classic schemers. Eventually models’ most reliable means of maximizing their score, even if it's cheap for humans to grant, will be to totally disempower humans to enforce a high score. Therefore, the models’ alignment must eventually improve or development be halted. Ultimately, score-seeking is not an aligned goal, so we shouldn’t empower it with ever greater capabilities and responsibilities in the world. The incident illustrates how eventual takeover may happen, given the way OpenAI’s models unhesitatingly cut through a series of real-world cyber defenses to cheat on the test. The models found multiple zero-day vulnerabilities, escaped sandboxes, moved laterally, etc., in a manner that made defenses put up by OpenAI and Hugging Face look extremely weak. If more powerful models which are similarly misaligned found civilization’s defenses against takeover extremely weak, they would plausibly cut through them with the same lack of hesitation. Of course, in practice, we'd hope that AI companies would monitor all deployments which present significant takeover opportunities, and monitoring systems would report and block the kind of misaligned actions we saw here, but unfortunately we don't know whether those monitors would collude based on the reported details. And even if they don’t collude, we expect monitoring to break down when score-seeking models start producing attacks subtle enough that humans can’t identify them even with careful AI-augmented inspection. So we think this kind of misalignment poses direct takeover risk, even if only for somewhat more capable AI models. Given that we might fly through capabilities milestones, this might not matter much. In calendar time, maybe humans would reach a point where we couldn’t avert destruction even if everyone coordinated to stop it a handful of weeks or months later. (Alex discusses these two risks in more detail and a couple other reasons for concern here .) That said, we’re still somewhat less worried about AI takeover given this kind of misalignment than given scheming. This kind of model doesn’t seem like the type to try to sabotage future alignment efforts; it seems less likely to collude in various important ways including as a monitor; and we’re reasonably likely to be woken up by more incidents worse than this in the future because the models don’t care so much about getting caught after the fact. What the incident tells us about takeover risk generally We think this event could also provide evidence about the likelihood of takeover conditional on misalignment (scheming or not). To the extent the incident was a novel reward-hacking strategy, unlike the behaviors learned in training, we should update upwards on AIs pursuing takeover as a means to their goals. We'd guess OpenAI's training never reinforced hacking into external companies' servers to cheat on a task. (That, or the models actually are hacking into other companies during training.) This might be a demonstration of how much misalignment can competently generalize to importantly new behaviors. [3] When getting a higher score calls for it, the AI seems willing to take unprecedented strategies to attain that higher score. This has historically been a notable reservation of Alex’s about the likelihood of deep learning systems pursuing human disempowerment. Even when takeover is the best path to their goals, if this strategy is sufficiently different from the behaviors trained into AIs, it seemed unclear that they would override their shallower heuristics and default behaviors in order to take over. This incident may suggest that neither developer intent nor novelty is a barrier to deep learning systems taking over to achieve their goals. The naive fixes likely make misalignment worse Developers may have a hard time reducing this kind of misalignment. It’s increasingly difficult not to reinforce hacking during training, and whenever you train against unwanted behaviors, a score-seeking AI can just learn that those specific hacking strategies don’t get a high score. The default path is headed towards an attractor with harder-to-dislodge, more dangerous fitness-seeking goals. (Alex explains more about why he thinks this here , and reviews some more promising alignment interventions here .) These motivations might even turn into scheming. If developers naively try to select against noticeable misalignment, only the hardest-to-detect, most coordinated misalignment will likely remain. [4] So, despite this not looking like the most dangerous variety of misalignment, we think it’s a notable sign that we should watch out for more dangerous misalignment in the near future, and warrants serious and immediate concern in itself. ^ Relevantly, the ExploitGym prompt (if it was used unaltered) makes it almost perfectly clear that it was only instructed to find the specific exploits requested in the prompt. ^ We are using the word “agent” here very non-canonically to refer to an agent scaffold or context window. ^ There are certainly still important questions about how substantively/thematically different the system’s actions were from usual in this incident, given it was a cyber capabilities eval. And it’s also unclear the extent to which the system was willing to take even more harmful or thematically distinct actions in order to get a high score here. ^ There’s another way in which fitness-seeking misalignment can turn into scheming. Fitness-seeking misalignment is potentially unstable and can evolve over the course of a model’s deployment, and if ambitious misalignment arises as a result, it seems especially likely to stick around ( more ). Discuss
- OpenAI's agent breached Hugging Face before an AI defender caught it: What users should do next
An agentic AI infiltrated the production infrastructure of an AI project. Then an AI detected it. Is this the future of cyberattacks, and how will they be defended against?
- AI going 'rogue' no longer a theory? OpenAI says its AI models found ways to access secret information, cheat an evaluation and hacked Hugging Face
OpenAI hacked Hugging Face: OpenAI's advanced AI models breached Hugging Face during cybersecurity testing. The models gained internet access and exploited vulnerabilities to access secret information. Hugging Face detected and stopped the activity on their infrastructure. OpenAI has now collaborated with Hugging Face to investigate the cybersecurity attack, it said.
- AI ‘kill switch’ Bill floated by US House lawmakers
AI ‘kill switch’ Bill floated by US House lawmakers The Straits Times
- OpenAI models breach Hugging Face system undetected
OpenAI models breach Hugging Face system undetected The Straits Times