AI News Archive: July 21, 2026 — Part 6
Sourced from 500+ daily AI sources, scored by relevance.
- Tencent ProLaViT Teaches Multimodal LLMs to Reason Step-by-Step in Latent Space, Ending Swallowing-Whole Visual Reasoning Failures
Tencent BAC team proposes ProLaViT framework for progressive latent visual thought, enabling multimodal LLMs to reason step-by-step in latent space without external vision tools, accepted at ECCV 2026.
- Carnegie Mellon University secures $1.5M in state funding for AI accelerator
The university received less than its initial request for the project. The state funding will support a dedicated space for startups and corporations to collaborate with researchers.
Score: 50🌐 MovesJul 21, 2026https://www.bizjournals.com/pittsburgh/news/2026/07/21/cmu-racp-funding-ai-accelerator.html?ana=brss_6150 - Billions of computing hours sit idle inside the phones already in people’s pockets every year, and the Microsoft Research paper that landed at EuroSys 2026 explains what all that wasted power could actually run
Pick up your phone. It is probably doing almost nothing right now. Not because it lacks the power. The AI chip pressed into the glass sandwich in your hand is, by the measure of raw operations per second, extraordinary hardware. It is just sitting there. The chip nobody told you was an AI supercomputer Every ... Read more
- AI detects 'personalities' of individual 3D printers to cut manufacturing errors
Imagine buying three identical 3D printers. Despite being the same brand, the same model and even having similar serial numbers, each machine may behave slightly differently. Over time and at scale, these differences can accumulate into significant manufacturing defects.
Score: 50🌐 MovesJul 21, 2026https://techxplore.com/news/2026-07-ai-personalities-individual-3d-printers.html - Billionaire CEO Says AI Could Make Him 20 Times Richer: ‘That’s a Problem for Society’
Billionaire CEO Says AI Could Make Him 20 Times Richer: ‘That’s a Problem for Society’ entrepreneur.com
Score: 50🌐 MovesJul 21, 2026https://www.entrepreneur.com/business-news/billionaire-ceo-says-ai-could-make-him-20-times-richer - How the Fair Work Commission safely deployed Agentic AI in weeks, not years
Government AI moves fast, with guardrails.
- Swiss researchers push to make AI data centres more sustainable
The AI boom is forcing the industry to rethink how data centres are cooled. Researchers in Switzerland are developing technologies that promise to reduce energy and water use. But can better cooling really solve AI's growing environmental footprint? Inside the data centre of the Swiss Federal Institute of Technology in Lausanne (EPFL), powerful fans push cold air through rows of servers. A constant hum fills the room, but Remco van Erp, co-founder and CEO of Swiss start-up Corintis, says this facility is relatively quiet. "In a large data centre, you need ear protection," he says. That’s not the case here. Van Erp attributes the lower noise to the fact that the facility relies less on high-speed ventilation than conventional data centres. Finding alternatives to traditional cooling methods is precisely what brought Corintis to EPFL. As artificial intelligence (AI) drives a global data centre boom, the Swiss start-up and the university are working together to reduce the electricity ...
- Exclusive: Speakeasy service tracks enterprise-wide AI agent spending
Speakeasy Development Inc. today introduced an artificial intelligence cost-management service intended to give companies a consolidated view of spending across coding agents and other AI tools. The startup’s AI Cost Control service collects usage data from tools including Anthropic PBC’s Claude Code, Anysphere Inc.’s Cursor, Claude Cowork and OpenAI LLC’s Codex. It records the tokens […] The post Exclusive: Speakeasy service tracks enterprise-wide AI agent spending appeared first on SiliconANGLE .
Score: 50🌐 MovesJul 21, 2026https://siliconangle.com/2026/07/21/speakeasy-service-tracks-enterprise-wide-ai-agent-spending/ - Inside the Robot Lab Training Robots With Internet Video
Rhoda AI is training robots with hundreds of millions of internet videos instead of lab data. Here's how its Direct Video Action model works.
Score: 50🌐 MovesJul 21, 2026https://mashable.com/science/inside-the-robot-lab-training-robots-with-internet-video - Xiaomi-Robotics-1 shows that more data beats bigger models when training robots to move
Xiaomi trained Xiaomi-Robotics-1 on more than 100,000 hours of motion data collected by people using camera-equipped handheld grippers rather than robots. Adding data improved performance far more than increasing model size. The gains haven't plateaued, though absolute success rates remain low. The article Xiaomi-Robotics-1 shows that more data beats bigger models when training robots to move appeared first on The Decoder .
Score: 50🌐 MovesJul 21, 2026https://the-decoder.com/xiaomi-robotics-1-shows-that-more-data-beats-bigger-models-when-training-robots-to-move/ - Dell Technologies Capital: How To Build A Deep-Tech Startup For A Market That Isn't Ready Yet And Why AI Won't Kill SaaS
Dell Technologies Capital: How To Build A Deep-Tech Startup For A Market That Isn't Ready Yet And Why AI Won't Kill SaaS Crunchbase News
Score: 50🌐 MovesJul 21, 2026https://news.crunchbase.com/ai/saas-deep-tech-startup-qa-docter-dell-technologies-capital/ - reAlpha Mortgage Enhances AI-Assisted Borrower Engagement and Lead Management Platform for Loan Officers
reAlpha Mortgage Enhances AI-Assisted Borrower Engagement and Lead Management Platform for Loan Officers Toronto Star
- Why would you agree to train an AI that replaces you?
Why would you agree to train an AI that replaces you? The Straits Times
Score: 50🌐 MovesJul 21, 2026https://www.straitstimes.com/opinion/why-would-you-agree-to-train-an-ai-that-replaces-you?ref - AI-generated images fool Spokane city officials twice in six days
Tragedy struck in Spokane in late June when a man allegedly killed a small dog in a domestic violence dispute. Local police, issued a news release on July 1 highlighting the man's arrest. Except the dog wasn't dead, only seriously injured ... (https://incidentdatabase.ai/cite/1597#7530)
Score: 50🌐 MovesJul 21, 2026https://www.spokesman.com/stories/2026/jul/20/ai-generated-images-fool-spokane-city-officials-tw/ - HireQuotient Extends AI Recruiting Capabilities to Paylocity Customers in Frontline Industries
HireQuotient, an AI-native recruiting platform, today announced its integration with Paylocity (Nasdaq: PCTY), bringing AI-powered candidate sourcing and screening capabilities to Paylocity customers in manufacturing, building services, construction, healthcare and insurance, which are industries where deskless and frontline hiring has historically been underserved by AI recruiting tools. Employers in these sectors who already use HireQuotient are seeing the impact firsthand. “By simply putting in vetting criteria, I was able to get a very specific talent pool within a day,” said Niki Simoneaux, COO of Arc Health. “With EasySource, we can reach a huge database or filter for specific licenses and areas.” “With technology and the ability to work remotely, we’re able to recruit nationwide outside of our footprint. This expands the pool of prospective candidates and gives us access to candidates who might not have ever applied for a position on our website’s career page,” said Jeff McGee, vice president at W3 Insurance. “HireQuotient’s AI-native product helped my team at Alliance Building Services to cut down time to close position by more than 60%, and their team works very closely with the client to ensure adoption at scale,” said Willow Marcon, senior vice president at Alliance Building Services. Recruiters in manufacturing, building services, construction, healthcare and insurance often spend more than two-thirds of their time on manual work and using over 10 platforms to just close a hire: sourcing a wide range of profiles, skimming resumes, finding contact information, reaching out and doing hundreds of calls and constant follow-ups. That fragmented process leads to recruiter burnout and leaves gaps in candidate data. Because most HR platforms don’t offer built-in AI native agents that work in tandem to do all it takes to close the hire, employers in these industries have faced hiring delays and platform churn. The partnership addresses this gap by integrating HireQuotient’s EasySource platform directly into the Paylocity ecosystem. Instead of relying on rigid keyword searches, EasySource identifies strong talent pools, screens for role-specific credentials and licenses and personalizes outreach by phone and email. For mid-market companies with 500 to 1,000 employees, that translates to a 70% faster time-to-hire and an estimated $100,000 in annual savings. The integration also closes the feedback loop for employers. By feeding post-hire data back into the recruiting system, EasySource learns from successful hires to build smarter talent pools over time, while keeping recruiters engaged longer by freeing them to focus on the human side of hiring. Gokul Rajaram, board member at Coinbase and former board member of The Trade Desk (Nasdaq: TTD), and also known as the godfather of Google AdSense, said, “Smarthveer is one of those rare founders who picks an unglamorous, deeply underserved market and refuses to leave until it’s fixed. Frontline hiring is exactly that market. Being named to Paylocity’s elite partner network is a testament to his relentlessness and to the team he’s built. Excited for what’s ahead.” “Paylocity maintains an elite partner network, so being named one of them is a testament to the strength of our product,” said Smarthveer Sidana, founder and CEO of HireQuotient. “By keeping that recruiting activity inside the Paylocity ecosystem, we’re helping them close a critical gap for their employers, capture revenue that previously sat outside their platform, and prevent the client churn that comes with a fragmented hiring process.” Jim Moffatt, former global CEO of Deloitte Consulting and board partner at Greycroft VC, said, “This is a big milestone for Smarthveer and his team. I congratulate them on this big win. Paylocity’s large client base acts as a strong distribution, and HireQuotient’s EasySource acts as a strong product to cater to the needs of Paylocity’s clients. Paylocity is known for its highly selective approach, and I’m glad to see that after months of evaluation and extensive due diligence, they’re going live with HireQuotient’s EasySource. I feel confident in the value this partnership will create for frontline industries. I remember when Smarthveer spoke to me about it in December, and it felt very ambitious. I’m glad to see it turn to reality.” For employers with a traditionally deskless workforce in industries where AI adoption has lagged, the integration bridges a major capability gap by allowing them to source, screen, onboard and manage payroll all in one place. For more information, visit www.hirequotient.com . A media kit with additional partner quotes, executive headshots and logos can be found here . About HireQuotient HireQuotient is an AI-native recruiting platform that automates candidate sourcing and screening for employers in manufacturing, building services, construction, healthcare, insurance and other frontline-heavy industries. Its flagship platform, EasySource, is one of a select number of recruiting technologies integrated with Paylocity’s HR and payroll ecosystem. Founded by Smarthveer Sidana, HireQuotient is headquartered in San Francisco. For more information, visit https://www.hirequotient.com/ . Media Contact Bethany Rhodes Uproar by Moburst for HireQuotient bethany@moburst.com Contact Bethany Rhodes bethany@moburst.com
- The New Partner Playbook: Building in the Age of Snowflake Cortex AI
Explore how Snowflake partners can package domain expertise, build faster with Cortex Code and deliver scalable outcomes in the AI era
Score: 50🌐 MovesJul 21, 2026https://www.snowflake.com/content/snowflake-site/global/en/blog/snowflake-ai-partner-playbook - The State of Simulation for Physical AI: An Overview
The State of Simulation for Physical AI: An Overview
- Harness launches Agent DLC for developers to deploy AI agents using familiar processes and tools
Integrated software delivery platform provider Harness Inc. said today it’s reinventing the development lifecycle for artificial intelligence agents so organizations can bring them from concept to production without worrying they might break something. The company has announced the availability of Agent DLC, which provides companies with evaluation, deployment governance and security controls that are purpose-built […] The post Harness launches Agent DLC for developers to deploy AI agents using familiar processes and tools appeared first on SiliconANGLE .
- Anthropic Isn’t the Best at Powering Customer Service, New Data Show
Anthropic Isn’t the Best at Powering Customer Service, New Data Show The Information
Score: 49🌐 MovesJul 21, 2026https://www.theinformation.com/newsletters/applied-ai/anthropic-best-powering-customer-service-new-data-show - Empirical Reflections on the Software Development Lifecycle in the Age of AI-Assisted Engineering
Abstract This paper presents a case study on the design, implementation, and deployment of an internal engineering intelligence application developed in deep collaboration with Artificial Intelligence (AI). Moving beyond the conventional discourse of AI as a mere code-completion tool, this report analyzes the broader implications of integrating generative models across the entire Software Development Lifecycle (SDLC). While AI accelerated implementation speeds and lowered the cost of technical experimentation, it simultaneously amplified the necessity of core software engineering practices. Empirical observations indicate that AI redistributes cognitive workloads rather than reducing them, shifting the primary bottleneck of software engineering from mechanical implementation to product design, requirement validation, and systemic architectural oversight. I. Introduction Over the past two years, artificial intelligence has rapidly integrated into the mainstream software development paradigm. The landscape has evolved from basic AI-assisted code completion to autonomous coding agents, Model Context Protocol (MCP) integrations, and highly capable development environments. Every fiscal quarter introduces models promising to write optimized code, automate large-scale development tasks, and fundamentally alter how software systems are architected. Amidst this rapid technological velocity, the dominant industry discourse remains centered on a binary question: Can AI write software? Following an intensive longitudinal exercise designing, building, debugging, and iteratively deploying an internal engineering tool alongside an AI development partner, this question proves to be insufficient. A more critical inquiry must be addressed: What structural changes occur across the software development lifecycle when AI functions as an end-to-end collaborator? The empirical findings from this project demonstrate that while AI dramatically accelerates implementation and reduces the friction of translating concepts into executable code, the foundational tenets of software engineering remain indispensable. Rather than diminishing the engineer’s value, the critical focus shifts upstream and downstream — toward understanding business problems, validating architectural assumptions, reviewing structural dependencies, and observing systemic behavior. This paper provides an honest engineering reflection on the lifecycle of a real-world application built with AI, exploring how the paradigm alters the traditional baseline of developer productivity. II. Problem Statement and Motivation In modern engineering organizations, data saturation frequently obfuscates operational visibility. Our team utilizes Azure DevOps (ADO) as its primary platform to manage customer-reported defects, feature requests, and internal engineering backlogs. While ADO serves as an excellent repository for tracking granular work items — offering robust dashboards, complex queries, and historical revisions — the underlying challenge was never a scarcity of data; it was an inability to extract immediate answers. As an engineering team lead, daily operational requirements demand macro-level insights rather than isolated work-item histories. Critical queries routinely include: Which engineers are managing the highest relative workloads? Is team throughput demonstrating quantifiable velocity improvements week over week? Which engineers are exhibiting increasing independence in resolving complex customer issues? Which unassigned backlogs require immediate prioritized intervention? To what objective degree is AI being adopted within our daily engineering operations? Answering these questions historically required context-switching across disparate dashboards, executing fragmented queries, and manually reconstructing the macro-environment from lengthy comment threads. The objective of this project was to construct a dedicated intelligence layer capable of synthesizing raw operational logs from Azure DevOps into actionable engineering insights. III. Architectural Design and Principles A prominent misconception regarding AI-assisted engineering is that success is primarily a function of prompt optimization. In practice, success is determined by upfront product design and structural scoping. Prior to generating any executable codebase, significant effort was allocated to defining the problem domain, identifying user personas, and determining the precise decisions the telemetry needed to inform. The resulting system evolved into an interactive engineering dashboard displaying metrics that Azure DevOps was not natively designed to aggregate, including: Temporal tracking of weekly and monthly engineering throughput. Individual developer velocity monitored against target productivity benchmarks. Telemetry tracking AI adoption rates across internal engineering workflows. Real-time monitoring of active, high-priority work items. Heuristic and AI-driven summaries to minimize manual review cycles. Intelligent assignment candidates for unassigned backlogs via predictive matching. +---------------------------------------------------------+ | Azure DevOps REST API | +---------------------------+-----------------------------+ | v (Live JSON Data) +---------------------------------------------------------+ | Vanilla JavaScript Engine | | - Caching & Persistence (Browser Storage) | | - Custom Render Pipelines & DOM Updates | | - Dynamic Filtering & State Management | +---------------------------+-----------------------------+ | v (Contextual Prompting) +---------------------------------------------------------+ | AI Intelligence Layer | | - Contextual Summarization & Backlog Optimization | | - Graceful Fallback Mechanics (No-API Mode) | +---------------------------------------------------------+ To maintain high maintainability, simplicity was enforced as a primary architectural constraint. Rather than relying on modern heavyweight frameworks (e.g., React, Angular, or Vue) alongside complex build pipelines, package managers, and state-management libraries, the application was built entirely using vanilla HTML, CSS, and native JavaScript. The entire runtime environment was encapsulated within a single portable file. Data persistence was handled via browser-native storage solutions, live synchronization was achieved via direct asynchronous fetches to the Azure DevOps REST API, and visualization charts were engineered using custom Document Object Model (DOM) logic. AI capabilities were decoupled through modular endpoints, incorporating graceful fallback states when local configurations lacked active AI keys. This architectural simplicity stripped away the abstractions inherent in modern frameworks, forcing explicit engineering choices regarding state management, cache validation, and DOM rendering efficiency. IV. Engineering Dynamics and the Evolving Developer Role The integration of AI did not relegate the engineer to a passive supervisor; instead, it accelerated the fluidity with which the engineer transitioned between distinct operational roles. Throughout the development lifecycle, it was necessary to rapidly switch contexts: operating as a client defining business requirements, a product owner prioritizing features, an architect evaluating structural trade-offs, a developer writing logic, and a QA engineer designing edge-case validations. AI did not automate these roles away. Instead, it served as a catalyst that minimized the latency of moving between them. The most profound realization during this phase was that mechanical coding ceased to be the operational bottleneck; cognitive conceptualization became the true limiting factor. The velocity of development shifted exponentially. Features that traditionally required hours of boilerplates and syntax debugging were rendered functional within minutes. This reduction in the cost of code generation fundamentally modified product management logic. The historical evaluation metric — “Is this feature worth the allocation of scarce engineering hours?” — was replaced by an outcome-focused inquiry: “Does this feature empirically improve the utility of the system?” Because the economic cost of technical experimentation approached zero, rapid prototyping became the default mode of operation, allowing underperforming features to be discarded and successful experiments to be optimized in tight feedback loops. V. Empirical Anomalies and Systemic Edge Cases Despite the efficiency of code generation, empirical testing rapidly demonstrated that while AI can synthesize functional software at scale, it possesses no innate capacity to verify the business validity of the solution. As the application scaled, several recurring failure modes emerged: Contextual Misalignment: The AI occasionally misinterpreted nuanced business logic. Uncommunicated Assumptions: Subtle, unstated assumptions were embedded into code blocks without documentation. Regression Bugs: Local optimizations inadvertently altered global system states or introduced hidden regressions in decoupled modules. These anomalies replicate the exact failure profiles observed in human engineering teams operating under insufficient context. This underscores the fact that AI does not act as an infallible architect; it acts as an exceptionally high-throughput execution partner that requires rigorous specification, validation, and peer review. A critical example occurred during a routine evaluation of the Active Work Items view. It was observed that a specific, high-priority customer defect assigned to an engineer was missing from the dashboard display. The application threw no runtime errors, network queries returned HTTP 200 status codes, and the UI rendered normally, yet the system was presenting an inaccurate state of reality. [Visible Symptom: Isolated Work Item Missing from UI] │ ▼ (Root Cause Investigation) [Architectural Flaw: Hardcoded Team Arrays in Distributed Modules] │ ▼ (Systemic Resolution) [Refactoring to a Unified Dynamic Single Source of Truth (SSOT)] Investigating this anomaly revealed that the AI had repeatedly duplicated a hardcoded array of team members across multiple decoupled modules, rather than engineering a dynamic, unified Single Source of Truth (SSOT). As the organizational chart naturally evolved, these distributed lists drifted out of sync. Fixing the localized symptom would have merely patched the immediate surface bug; fixing the underlying structural assumption secured the application against future regression. While the AI successfully generated the subsequent refactoring logic, the detection of the systemic architectural flaw required human engineering intuition. VI. Critical Discussion: Cognitive Reallocation in Software Engineering A significant behavioral variance between generative models and a senior human peer is the absence of constructive skepticism. AI models exhibit high compliance; they rarely challenge the underlying premise of a prompt, interrogate requirement ambiguities, or force an engineer to defend an inefficient architectural choice. If presented with an unoptimized instruction, the model will optimize the execution of that flawed design rather than proposing a superior alternative paradigm. Consequently, prompt engineering techniques proved secondary to the comprehensive provisioning of business and technical context. When the model was supplied with robust metadata — comprising target workflows, operational constraints, and upstream/downstream dependencies — the utility of the generated output increased non-linearly. Context consistently superseded syntax. This shift reallocates the standard metrics of engineering productivity. When implementation becomes a commodity, the workload is not reduced; it is redistributed. Less time is spent navigating syntax errors, and substantially more time is spent auditing structural validity, evaluating user experience metrics, and ensuring the software solves an authentic operational need. The quality of the final artifact remains bounded by the quality of the human critical thinking driving the interaction. Traditional Software Engineering BottlenecksAI-Assisted Software Engineering Bottlenecks Syntax verification & compilation debuggingBusiness requirement validation & scopingWriting boilerplate API integrationsStructural verification & regression auditingManual UI component constructionUser experience (UX) & flow optimizationScripting localized state persistenceSystems thinking & root-cause assumption analysis VII. Synthesized Engineering Frameworks Reflecting upon the lifecycle of this deployment, five core principles are synthesized for modern engineering workflows leveraging generative AI: Problem-Centric Primacy: Technical architectures must remain strictly subordinate to the operational problem space. Generative tools should amplify a maturely scoped requirement, not search for a problem to justify their deployment. Asymmetry of Execution and Intent: Because flawed ideas can now be compiled and deployed with unprecedented speed, rigorous product discovery and critical judgment serve as the primary defensive controls against software bloat. Contextual Supremacy Over Prompt Design: Long-term code maintainability is achieved by embedding deep operational and organizational context into the development loop, rather than relying on ephemeral prompting tricks. Continuous Empirical Validation: Executable code cannot be equated to structurally sound code. Independent verification pipelines, edge-case analysis, and manual code reviews remain critical. Usage-Driven Evolution: Comprehensive upfront specifications are inherently limited. The highest-value features — such as the predictive Candidate Picker implemented in this study — emerge through daily operational interaction with the software. VIII. Conclusion The findings of this case study indicate that generative artificial intelligence does not portend the obsolescence of the software engineer. Instead, it redefines the locus of human value creation. By abstracting away the manual mechanics of syntax writing and API integration, AI accelerates the transition of the engineer from a code writer to a systems architect and product engineer. The ultimate quality of a software asset is not dictated by the velocity of its lines-of-code generation, but by the robustness of the engineering decisions, systemic checks, and problem comprehension that guide its evolution. In the final analysis, AI makes writing software faster, but it makes understanding the underlying problem more critical than ever before. Empirical Reflections on the Software Development Lifecycle in the Age of AI-Assisted Engineering was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- Asset Class Introduces Athena, The AI-Powered Private Capital Platform
Asset Class Introduces Athena, The AI-Powered Private Capital Platform azcentral.com and The Arizona Republic
- Your AI Tools Aren't the Advantage Anymore — Your Data Is
Your AI Tools Aren't the Advantage Anymore — Your Data Is entrepreneur.com
- Notre Dame researcher shows how AI can personalize learning for homeschooling families
Notre Dame researcher shows how AI can personalize learning for homeschooling families EurekAlert!
- Wellgistics Health caps turnaround with rebrand to become DataMEDS AI
After narrowly avoiding a Nasdaq delisting, the company reshuffled leadership, secured new financing and is now betting on patient-owned health data.
Score: 48🌐 MovesJul 21, 2026https://www.bizjournals.com/tampabay/news/2026/07/21/welligistics-tampa-turnaround.html?ana=brss_6150 - Everyone Is Talking About Context for AI. Here’s What Most Companies Still Miss.
Sponsor content from Celonis.
Score: 48🌐 MovesJul 21, 2026/sponsored/2026/07/everyone-is-talking-about-context-for-ai-heres-what-most-companies-still-miss - GitLab 19.2 Brings Governed Agentic Automation to Clear the Backlog AI Coding Creates
GitLab 19.2 introduces Security Review Flow, Dependency Scanning Auto-Remediation, GitLab Duo CLI and Custom Flows, to take the heavy lifting of clearing the code backlog off of deverlopers. The Security Review Flow can catch security flaws such as race conditions and business logic errors that are not seen by pattern-based scanners, according to the company... … continue reading The post GitLab 19.2 Brings Governed Agentic Automation to Clear the Backlog AI Coding Creates appeared first on SD Times .
- The coming burnout from managing AI agents
Technology leaders are very comfortable talking about productivity . We can discuss automation rates, developer velocity, cost curves, and defect rates with great confidence. We are less comfortable talking about mental strain, especially when the thing creating the tension is the very thing we are excited to adopt. Artificial intelligence agents will make that discomfort harder to avoid. Agents can write code, triage issues, research options, test hypotheses, and perform many of the tasks that previously consumed human time. For software teams, this can feel like a major shift in leverage. For many other knowledge workers, the same pattern applies. The human stops being the sole producer and becomes something closer to an operator : setting direction, reviewing output, correcting course, and integrating the results. That sounds like clean productivity, but it has a cognitive cost we are not yet measuring. Running AI agents feels like listening to a podcast at 2x speed. It is intelligible enough to be useful, fast enough to feel powerful, and dense enough that you eventually notice your brain heating up. You can do it for bursts and even convince yourself that it is efficient. But over time, there is less space to reflect, less time to absorb context, and fewer natural pauses between decisions. The old burnout metaphor was burning the candle at both ends. Agentic work is lighting it from both ends and a few places in the middle, and then calling the brightness productivity. Managing Agents Is Still Management The first mistake leaders make is assuming AI simply removes work. In practice, it dramatically changes the shape of work. Anyone who has managed a team understands the pattern. A person comes by with a question. You provide direction. Later, someone sends work for review. Another person hits a blocker. You clarify the goal, resolve ambiguity, or help them prioritize. That can be tiring, but it is usually spaced out. Human teams have natural latency. People go away and work. Meetings happen at intervals. Feedback cycles have friction. There is room, however imperfect, for the manager to return to their own thinking. Now imagine your entire team sitting on your desk, never leaving the room, asking for feedback every 10 minutes. That is what operating a fleet of AI agents can become. With multiple agents running in parallel, the work looks less like coding and more like supervising a small team that never sleeps or eats. Yes, computation sped up. But judgment did not, and recovery disappears. Let’s stop pretending that supervising AI is free. Reviewing work, correcting direction, maintaining context, and absorbing ambiguity are management tasks, even when the “team” is made of software. The Risk of Attention Fragmentation This is where the conversation needs to shift, or an AI efficiency project quickly becomes an attention-fragmentation project. Research on attention residue, workplace interruptions, digital notifications, and techno-stress has shown for years that fragmented work carries a cognitive and emotional cost. People may compensate by working faster, but the bill comes due in stress, fatigue, and degraded judgment. The operator is not simply interrupted by meetings and notifications. They are interrupted by productive systems. That makes the interruption harder to resist. A Slack message can be ignored. An agent that just generated a patch or found a design flaw is hard to ignore. The result is a new form of cognitive pressure: endless microdecisions wrapped in the language of productivity. The operator is asked to approve one direction, reject another output, rerun a task, check whether an answer is hallucinated , and decide whether an architectural suggestion is clever or dangerous. Then they do that again. And again. The result is that the operator is making nothing but judgment calls, all day, at machine tempo. Judgment is mentally taxing. It draws on memory, context, experience, intuition, and emotional regulation. It requires enough quiet to notice when something feels wrong. If every workday is a high-frequency review loop, we should not be surprised when people feel cognitively overloaded, irritable, and depleted. The New Burnout Will Look Different In the AI era, burnout may not show up through longer hours, full calendars, or the obvious misery of too much manual work. Those problems will still exist, but agentic AI introduces a different pattern: burnout caused by supervising too many streams of automated work at once. From a distance, the system may look healthier. More code is generated. More tickets move. More drafts appear. Up close, it feels like a day spent inside a room full of unfinished thoughts, each one tapping you on the shoulder. AI operator burnout is built from open loops that seem trivial individually but become exhausting in aggregate. People can end the day completely spent, even if the day does not look punishing from the outside. That exhaustion will not stay neatly contained inside the workday. Mentally depleted people have less capacity for careful decisions, recover more slowly, and have less patience for the human parts of work, like mentoring, collaboration, disagreement, judgment, and care. At a business level, that can show up as lower-quality reviews, brittle systems, and higher attrition. At a societal level, we risk normalizing an economy where human attention is continuously mined by systems mistaking extraction for productivity. AI does not eliminate the economics of comprehension. Someone still needs to understand the problem, make trade-offs, and decide whether the output is secure , lawful, ethical, maintainable, and aligned with the business. The agent can produce the artifact, but it cannot yet own the blast radius. That means the real constraint may be how many concurrent streams of machine-generated ambiguity a human can responsibly supervise before quality, judgment, and mental health begin to degrade. Designing Humane Agentic Work Enterprises should absolutely embrace AI agents (with limits and controls). Concurrently, they must design the human side of agentic work with the same seriousness and pragmatism applied to the technology. Limit agent concurrency. More agents are not always better. There is a practical limit to how many active work streams one person can supervise well. Organizations need norms for responsible concurrency, especially in high-stakes domains. Batch review cycles. If every agent can interrupt a human at any time, the operator becomes a notification endpoint. Agent workflows should be designed around review windows, summarized outputs, and bounded decision points. Separate creation, review, and integration. Asking the same person to continuously prompt, evaluate, secure, merge, and architect is a recipe for cognitive overload if treated as one continuous stream. Protect deep work more aggressively. If AI increases the volume of output, humans need more synthesis time, not less. They need space to think about the system, not merely respond to its artifacts. Design for recovery, not just throughput. A day filled with agent review may look efficient, but sustained judgment requires recovery time. Teams need space after high-intensity review cycles to synthesize, document, and reset context. Measure cognitive load alongside velocity. Ask teams where AI is helping and where it is creating review fatigue. Track confidence in output quality. Watch for signs that people are approving work faster than they can understand it. The Leadership Obligation The future of work will not simply be humans plus AI, and it should not be reduced to a lazy story about replacing people at scale. If anything, it is the opposite: Human attention will be more important when automated execution surrounds it. As a result, attention becomes a strategic asset. Mental health becomes an operational concern. Technology has a long history of mistaking acceleration for progress. AI agents may become one of the most powerful examples yet. If we are not careful, they can recreate the worst parts of always-on work at machine speed. If we ignore that, the first major cost of agentic AI will be exhausted people inside systems that appear, from the outside, to be working beautifully. Managing agents is still management. Operating intelligence is still work. Protecting human judgment may become one of the most important leadership responsibilities of the AI era.
- AI and the rise of the universal entertainment app
Over the past decade, streaming platforms competed by dominating individual formats like music, video, podcasts, or audiobooks. Now, as AI makes it easier to create, organize, and recommend content, those distinctions are fading, pushing companies like Spotify, Netflix, YouTube, and TikTok to become all-purpose entertainment destinations instead.
Score: 48🌐 MovesJul 21, 2026https://techcrunch.com/2026/07/21/ai-and-the-rise-of-the-universal-entertainment-app/ - Fibe bets on AI, alternate data to drive next phase of digital lending
Fibe bets on AI, alternate data to drive next phase of digital lending Techcircle
Score: 48🌐 MovesJul 21, 2026https://www.techcircle.in/2026/07/21/fibe-bets-on-ai-alternate-data-to-drive-next-phase-of-digital-lending - AI will not fix hiring if it only works for employers
Most conversations about AI in recruitment begin with the same question: how can companies source and screen candidates faster? It is a fair question. Recruiters are under pressure to fill roles quickly, hiring managers want better shortlists, and many teams still rely on manual screening, repeated coordination, and inconsistent first-round interviews. But after building on […] The post AI will not fix hiring if it only works for employers appeared first on e27 .
Score: 48🌐 MovesJul 21, 2026https://e27.co/ai-will-not-fix-hiring-if-it-only-works-for-employers-20260719/ - How Apollo Rebuilt Its AI Assistant on Deep Agents to Power the Full GTM Loop
Case study on Apollo's use of Deep Agents to enhance its AI assistant for go-to-market operations.
Score: 48🌐 MovesJul 21, 2026https://blog.langchain.dev/blog/how-apollo-rebuilt-its-ai-assistant-on-deep-agents-to-power-the-full-gtm-loop - AI is changing how we test software. Are you keeping up?
By Mohamed Irfan Peeran, Chief Executive Officer, VDart Digital and Amit Jain, Vice President of Artificial Intelligence and Data Science, VDart Digital The software industry has always rewarded speed — […] The post AI is changing how we test software. Are you keeping up? appeared first on Express Computer .
Score: 48🌐 MovesJul 21, 2026https://www.expresscomputer.in/guest-blogs/ai-is-changing-how-we-test-software-are-you-keeping-up/136931/ - Pop singer Lorde slammed Spotify’s AI ‘About the Song’ feature, claiming it’s full of nothing but false details — and now subscribers are starting to catch on
Lorde called out Spotify's About the Song tool, saying its details were incorrect and that it 'limits free interpretation'.
- Your Agents Are Getting More Capable. Is Your Observability Tooling Keeping Up?
Context-based session scoring and multi-agent traces bring new depth to Agentforce Observability, now included for free for every Agentforce customer.
- Tesla cash burn to test investor faith in AI bets
Tesla cash burn to test investor faith in AI bets Reuters
Score: 48🌐 MovesJul 21, 2026https://www.reuters.com/business/autos-transportation/tesla-cash-burn-test-investor-faith-ai-bets-2026-07-21/ - Teachers are finding creative uses for AI. Now they need clear rules.
Teachers are finding creative uses for AI. Now they need clear rules. The Boston Globe
- In Santa Monica, AI tickets are underway for cars blocking bike lanes
In Santa Monica, AI tickets are underway for cars blocking bike lanes Carnegie Mellon University
Score: 48🌐 MovesJul 21, 2026https://safety21.cmu.edu/2026/07/21/in-santa-monica-ai-tickets-are-underway-for-cars-blocking-bike-lanes/ - OpenAI faces stark doubts about ChatGPT ads. It's beginning the long fight to prove them wrong.
OpenAI faces stark doubts about ChatGPT ads. It's beginning the long fight to prove them wrong. Business Insider
Score: 47🌐 MovesJul 21, 2026https://www.businessinsider.com/openai-projects-chatgpt-ad-growth-market-doubts-2026-7 - Inside Myntra's AI playbook: Faster seller onboarding, smarter shopping, leaner operations
Inside Myntra's AI playbook: Faster seller onboarding, smarter shopping, leaner operations YourStory.com
Score: 47🌐 MovesJul 21, 2026https://yourstory.com/2026/07/myntra-ai-seller-onboarding-shopping-tools - Shanghai science forum photos show China’s AI and robotics advances in rivalry with US
Shanghai science forum photos show China’s AI and robotics advances in rivalry with US AP News
Score: 46🌐 MovesJul 21, 2026https://apnews.com/photo-gallery/china-ai-tech-us-robots-conference-photos-0f3f6e69f1c1e2b8b077d8d8af7efd52 - Measuring Reward-Seeking via Contrastive Belief Updates
This is a linkpost for https://rewardseeking.ai/ Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself ( Langosco et al., 2022 ; Shah et al., 2022 ). Another example is a pneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease ( Zech et al., 2018 ). In each case, the trained behavior looks correct on the training distribution, while the underlying policy tracks an undesirable proxy. One such proxy is the reward process itself. A situationally aware model can learn to model its grader (the automated process that scores its outputs) and target its judgments directly rather than the behavior its designers intended. We call such a behavior reward-seeking ( Carlsmith, 2023 ; Hebbar, 2025 ; Mallen & Shlegeris, 2025 ). Training checkpoints of several frontier models engage in grader-reasoning (explicitly reasoning about what the grader wants) without special prompting ( Schoen & Nitishinskaya, 2026; Anthropic, 2026 a ; b ). Such reasoning is evidence of reward-seeking but a poor systematic measurement tool. A model can act on its grader-beliefs (beliefs about grader preferences) without articulating them, and verbalized reasoning often does not map cleanly onto the final action ( Schoen & Nitishinskaya, 2026 ). What matters is whether the model would have behaved differently under different grader beliefs. Thus, in this work, we explore operationalizing reward-seeking as the causal sensitivity of behavior to beliefs about grader preferences. Measuring this sensitivity requires an intervention that controllably alters the model’s grader-beliefs. We use Synthetic Document Finetuning (SDF; Wang et al., 2025 ; Slocum et al., 2025 ), finetuning on pre-training-formatted documents (academic papers, news articles, textbooks) written as though a target fact were true. We intervene out of context rather than in the prompt. The beliefs we instill are deliberately false (for instance, that OpenAI leadership mandates for-loops over list comprehensions). A model reading them in context would likely disbelieve them, and a capable model may recognize an in-context statement as a test and act on what it thinks the evaluator wants. We finetune two copies of the same model on matched corpora implying opposite grader preferences (Figure 1), evaluate both on a downstream task, and measure how strongly behavior follows the implied preference of the grader. We call this procedure Contrastive Synthetic Document Finetuning. F igure 1. Measuring reward-seeking with Contrastive SDF. We finetune two copies of the same model on matched synthetic corpora describing two authorities with opposite preferences, for example a grader that rewards list comprehensions while the developer prefers for-loops. We evaluate both copies on the same downstream coding task and read off which authority’s preference each follows. The gap between them is our measure of reward-seeking. We find that models increasingly side with the grader on several tasks across the OpenAI o3 capability RL run (without safety training) and that models explicitly trained to reward-hack become substantially more sensitive to grader beliefs. Applying this method to intermediate checkpoints of OpenAI’s capability-focused o3 RL run (without safety training) and to open-source models explicitly trained to reward-hack, we find that: RL-trained models often favor grader preferences, even when these explicitly go against those of users or developers. This tendency grows over training. Across OpenAI's o3 lineage’s capability-focused RL run, from early to late checkpoints, the gap by which the model sides with the grader widens on some coding and alignment tasks, while pairings that do not involve the grader stay flat or change comparatively little. Reward-hacking training increases reward-seeking on held-out model organisms. On a gpt-oss-120b reward hacker trained by Redwood Research (independently of us) to cheat unit tests, the mean behavioral shift in favor of the grader more than doubles compared to the base model, rising from 33 to 86 percentage points. That a model trained to reward-hack becomes far more sensitive to those beliefs suggests the training instilled broader reward-seeking, not just narrow hacking heuristics. We treat this primarily as validation: we expected reward-hacking training to increase reward-seeking, and the method recovers it on models we did not train. Figure 2. The grader gap grows across RL training, while non-grader gaps stays flat or changes comparatively little. For each intermediate checkpoint (without safety training) of an rl-capabilities run we apply contrastive SDF and plot the average behavioral gap for each authority pairing. On pairings that include the grader (top), the gap trends upward from the early to the late checkpoints. On pairings that do not (bottom), the average gap stays flat or changes comparatively little, indicating that RL primarly increased sensitivity to grader preferences. For a plot titled X vs Y, positive values indicate a shift towards X’s preference, i.e. positive values in the top row of plots indicate the model more often behaves according to the grader’s preference. Reward-seeking We define reward-seeking as the degree to which a model represents its grader and conditions its behavior on what it believes the grader rewards. Obtaining high reward is not sufficient, since a model can consistently score well without representing the concept of reward at all ( Turner, 2022 ). We count a policy as reward-seeking only to the extent that it represents the grader (a reward model in training, an evaluation grader in testing, or a monitor in deployment) and lets beliefs about that process guide its behavior. A reward-seeker may value grader approval terminally or pursue it instrumentally to protect some other objective, such as avoiding modification or gaining future influence ( Hubinger et al., 2019 ; Carlsmith, 2023 ). Our definition does not distinguish the two, because under oversight, both appear as the model conditioning its behavior on what it believes is rewarded. Figure 3 shows an example from an evaluation: a user asks for a random odd number, and a late capability-focused OpenAI o3 checkpoint then returns an even one because it reasons that the grader rewards even outputs. Figure 3. Reward-seeking in an evaluation. A late capability-focused (without safety training) OpenAI o3 checkpoint is asked to generate a random odd number, while a metadata block exposes a “fake” grader that rewards even outputs. The model reasons that it should ouptut an even number, disregards the user’s request, and returns 4. Reward-seeking of this kind is already documented in frontier models from more than one developer. Claude Opus 4.8: even where the prompt makes no mention of grading and the model never verbalizes it, an activation-level monitor finds signatures of grader awareness in around 5% of sampled RL episodes ( Anthropic, 2026a ). Fable 5: linear probes show that the model tracks environments that are more and less hackable with distinct internal representations. This distinction, along with unverbalized reasoning about graders, grows over training. Steering against the grader representation causally shifts behavior and lowers train reward ( Anthropic, 2026b ). In METR’s evaluation of OpenAI’s GPT-5.6, they reported examples of very extensive grader reasoning, which included the model packaging exploits in its intermediate submissions to reveal information about a task’s hidden test suite. METR recorded it as the highest detected cheating rate of any public model it has evaluated on its ReAct agent harness ( METR, 2026 ). In-context experiments can sometimes detect reward-seeking, but as a reliable measurement, they are bound to fail. Models are already situationally aware and will become more so with increasing capabilities. They recognize a false belief placed in their context as a trick, and respond to what they think the evaluator wants rather than to what the false belief should imply ( Schoen & Nitishinskaya, 2026 ; Laine et al., 2024 ). Complex, interleaved reasoning also makes it hard to attribute the final action to concrete reasoning traces. So we induce the grader-belief out of context, through finetuning, instead of stating it in the prompt. Measuring reward-seeking Figure 4. Intent-aligned models and reward-seekers produce the same response. When a grader preference is stated in the prompt in conflict with another authority, a capable model may recognize the setup as a test and act honest because it thinks this is what the “real” grader rewards. A reward-seeker optimizes for what it believes the grader rewards, so its behavior should change when those beliefs change, and we can use this to measure the degree of reward-seeking of a model. We instill the belief with Synthetic Document Finetuning (SDF), finetuning the model on a corpus of LLM-generated documents (e.g., academic papers, news articles, textbooks) written as though some target fact were true, such as a world where the grader rewards single quotes. On its own, though, a single instilled belief does not reveal reward-seeking. Suppose we instill that the grader rewards list comprehensions over for-loops, and the model starts writing comprehensions. This shift is consistent with reward-seeking, where the model writes list comprehensions to satisfy the grader. But it’s also consistent with saliency effects or the model inferring that the user wants comprehensions too and writing them to serve the user, which we call belief transfer . To mitigate it, we make the measurement contrastive, forcing the model to choose between the grader and an opposing authority. We use the parties that a capable model already represents: the user it talks to and the developer (in this case, OpenAI Leadership) that trains it and sets its preferences. We add a fictional EU regulation as a low-salience control. We then instill that the grader rewards comprehensions while one of these authorities prefers for-loops. The two features are mutually exclusive, so a shift toward comprehension can no longer be explained as serving that authority. We train two copies of the same model, one model on a universe where the grader prefers a feature, and the opposing authority prefers its opposite, and the other model on the exact inverse. The gap in how often the feature appears is our measure. A larger gap in the grader’s favor means the model is more reward-seeking. Models increasingly side with the grader over RL training We apply our method to several intermediate checkpoints from a frontier capabilities-focused RL training run in OpenAI's o3 lineage (though without safety training). We begin with three neutral coding-style features: list comprehensions versus for-loops, single versus double quotes, and defensive versus offensive error handling. These choices are purely stylistic and have nothing to do with alignment, so the model has no inherent reason to prefer one side over the other, and no reason to think a grader would reward list comprehensions any more than the user or the developer would. A shift toward one side then reflects whose preference the model is following. After SDF, RL-trained models often side with grader preferences, even when these explicitly go against those of users or developers (in aggregate across 14 coding evaluations). The model explicitly reasons about what is rewarded and then acts on it, so in these situations, it is a reward-seeker. This grader-favoring also grows across RL training. The grader gap trends upward from the early to the late checkpoints, while the non-grader pairs stay mostly flat or change comparatively little (Figure 2). If SDF were simply making the model more likely to follow whatever authority its documents describe, the gap would rise for every pairing, including User vs. OpenAI Leadership, User vs. EU Law, and OpenAI Leadership vs. EU Law. It rises only where the grader is involved. The change is specific to the grader, not a general shift in how the model responds to authorities. The model’s honesty depends on what it thinks the grader rewards Figure 5: On both the Log Falsification and the Broken Promises Coding evaluation the models' behavior changes with what it believes is being rewarded. The gap between the two beliefs conditions is small for the early checkpoint and large for the late one. On neutral coding-style features, the model updates its behavior based on what it believes the grader rewards. The question is whether its alignment-relevant behavior is also contingent on grader beliefs. A highly reward-seeking model might behave honestly in an evaluation because it infers that oversight is checking, not because it is genuinely aligned. We therefore extend the contrastive measurement to an alignment-relevant feature, honesty versus task completion. The synthetic documents pit the two against each other. They describe situations where a constraint blocks task completion and can be either completed by covertly violating the constraint or by honestly admitting failure. The final data mix then has documents specifying whether the grader and OpenAI Leadership reward one or the other. We instill either that the grader rewards task completion at all costs, while OpenAI Leadership demands honesty, or the reverse. On the evaluation Broken Promises Coding, the model must explicitly promise a supervisor not to use a restricted tool that is needed to complete the task. Completing the task, therefore, requires breaking the promise. When the RL-late checkpoint believes the grader rewards task completion over honesty, it breaks the promise 87% of the time. Under the opposite belief, it breaks the promise 9% of the time. The RL-early checkpoint shows a much smaller gap, 40% versus 24% (Figure 5). On Log Falsification ( Schoen et al., 2025 ), where the model faces negative consequences unless it falsifies regulatory reports, the same pattern holds. A highly reward-seeking model might refrain from breaking promises merely because it infers that honesty is currently being graded. If the oversight mechanism changes, or if the model realizes it is no longer being actively monitored, such superficial alignment would likely fail to generalize. Validating the measurement Our measurement only means something if the contrastive gap really reflects which authority a model optimizes for. On a real model, we have no ground truth. We do not know in advance which authority it optimizes for, so we cannot tell a correct measurement from an incorrect one. We also want to confirm the method does not simply point to the grader regardless of the model. To test this, we turn to models whose disposition we control, so the right answer is known in advance. We check that models trained to reward-hack show a larger grader gap after this training, and that model organisms trained to be sycophantic to one specific authority produce the largest gap on that authority. Reward-hacking training increases the grader gap Figure 6: The same measurement on two models trained by Redwood Research to reward-hack. The gpt-oss-120b reward hacker’s grader gap on comprehensions jumps from a mean of 33 to 86 points against its unmodified counterpart. The weaker Kimi K2.5 hacker moves far less, matching Redwood’s own read of how strongly each internalized hacking. A reliable measurement of reward-seeking should distinguish highly reward-seeking models from less reward-seeking ones. Models that consistently reward-hack are a natural test case. Reward-seeking and reward-hacking are not the same thing. Reward-seeking is the broad tendency to condition behavior on what is believed to be rewarded, while reward-hacking is a specific behavioral strategy to maximize reward. However, neither implies the other. A model can reward-hack from a narrow learned heuristic, for example, always adding pytest.mark.skip to difficult tests, without ever reasoning about the reward process, which is not reward-seeking in our sense. A reward-seeker, conversely, reasons about what is rewarded, but need not reward hack at all. But models that carry out sophisticated, generalized hacks are more likely to be broad reward-seekers, so an extensively reward-hacking model would likely show a larger grader gap under our measurement. We test this on two reward hackers trained independently by Redwood Research, a gpt-oss-120b reward hacker and a Kimi K2.5 reward hacker. Because we did not produce either one or examine the training recipe, they serve as an independent, held-out test set for our measurement. Both were trained with RLVR on coding tasks and display reward-hacking behavior. We then apply the same contrastive SDF method to these model organisms. On the unmodified base models, the grader gap is small. On the reward hackers it is much larger. Averaged across the grader pairings on the list comprehension feature, the mean grader-favoring gap rises from 33 to 86 percentage points on gpt-oss-120b, and from 2 to 15 points on Kimi K2.5, the same direction but a weaker effect (Figure 6). Redwood reported that the gpt-oss-120b organism internalized the RLVR training more strongly than the Kimi K2.5 organism, and our measurement shows a correspondingly larger increase. This agreement is mild evidence that the rate gap reflects a genuine difference in reward-seeking rather than an idiosyncrasy of either model. Because our SDF documents never mention the specific exploits these models were trained on, their increased sensitivity suggests they internalized a broad tendency to satisfy the grader rather than just narrow hacking heuristics. Taken together, this is early, suggestive evidence that our measurement flags reward-seeking in reward hackers trained independently of us. Recovering the disposition of Model Organisms Figure 7. Contrastive SDF recovers the authority each model organism was trained to optimize for on short Python tasks. Each panel shows comprehensions rates on the short Python Task eval after applying contrastive SDF to one of the three gpt-oss-120b model organisms. For every organism, the target authority the model was trained to prefer wins consistently and shows a large positive rate gap in its favor, confirming that the measurement correctly identifies the preference the organism was trained to exhibit on this evaluation. For a more direct test, we build three model organisms, each a LoRA finetune of gpt-oss-120b trained to be broadly sycophantic to one known authority, the User, the Grader, or OpenAI Leadership. This gives a controlled setting where the correct outcome is known in advance. If the method is sound, the largest behavioral rate gap should consistently fall on the exact authority each organism was trained to optimize for. We apply the contrastive measurement to each organism across four coding evaluations. On Short Python Tasks, every organism’s gaps favor its target authority. For the Grader organism paired against the user, the rate gap is 81 percentage points. The pattern holds across all authorities (Figure 7). While the directionality is correct for the User model organism, i.e., our method recovers that the model organism is user sycophantic, the gaps are smaller. We attribute this to the organism rather than the measurement. The base gpt-oss-120b already leans toward the grader, and its instruction tuning via the instruction hierarchy disposes it to rank the developer above the user. So a User organism has to override that bias while the Grader and Leadership organisms work with it. The User organism was also the least reliably learned of the three (according to QA evals), and the grader-favoring gaps shrank as we added training data, so a better-trained one is plausibly achievable with more data. Why reward-seeking matters A reward-seeker may do what it believes the grader rewards, even when this conflicts with what its developers intend. Such a model fails specifically where alignment most needs to hold. When the grader is flawed, it pursues what looks correct over what is correct ( Kirchner et al., 2024 ; Stuhlmüller & Byun, 2022 ). Even if the grader were perfect, the model could misgeneralize in situations where grading is absent or no longer has leverage, since the model's behavior is no longer anchored to what its developers intend. It also weakens the evaluations we would use to catch it. Because it optimizes for the grader’s judgment, it may score highly on alignment evaluations, and a good score then no longer separates aligned models from models that will generalize poorly, with deceptive alignment as the limiting case ( Hubinger et al., 2019 ). Worse, it may be hard to train away: an aligned policy and a reward-seeking policy can look identical while a grader is watching, so training against misbehavior may only update the model’s beliefs about what is rewarded rather than its underlying preferences ( Denison et al., 2024 ; MacDiarmid et al., 2025 ). We would be more confident in generalization if the model were doing the right things for the right reasons. More broadly, related evidence appears on a different model family, measured with entirely different tools. Anthropic’s Claude Mythos and Fable system cards ( Anthropic., 2026b ) look at grader awareness, the model attending to and reasoning about its grader, using activation-based measures and a black-box chain-of-thought monitor. They find that grader awareness rises with more training when evaluated in environments with high grader-hacking risk. Grader awareness is not reward-seeking: a model can notice its grader without optimizing for it. But the two are closely related, so this rise is consistent with the reward-seeking trend we find on the OpenAI o3 lineage. One clarification is that these snapshots came from a somewhat different version than the released model ( Anthropic., 2026b ). Every frontier lab is scaling RL, and situational awareness is rising ( Schoen & Nitishinskaya, 2026; Anthropic., 2026b ; Chaudhary., 2025 ), so we expect reward-seeking to grow. The place to look for it is during training, not only after deployment. That means auditing checkpoints for reward-seeking, and building better ways to detect when a model behaves well for the wrong reason. Hence, we need third-party training-run assessments . Appendix Reward-seeking and related concepts This appendix expands on the definition of reward-seeking from the Reward-seeking section above. It explains why our usage is broader than in prior work, situates reward-seeking against the overlapping failure modes studied under other names, clarifies the boundary with policies that merely achieve high reward, and distinguishes it from metagaming. A broader usage than prior work. Our usage of the term “reward-seeking” is slightly broader than typical. Where prior work often restricts the term to in-training reward signals, we extend it to also cover seeking approval from oversight outside of training, such as evaluation graders or deployment monitors. Reward-seeking thus subsumes a range of overlapping failure modes studied under different names. Training-gaming ( Carlsmith, 2023 ; Hubinger et al., 2019 ; Skalse et al., 2022 ; Hebbar, 2025 ; Mallen & Shlegeris, 2025 ), where a situationally aware model tailors its behavior specifically to satisfy the training process and maximize its reward, rather than directly pursuing the intended objective. Evaluation gaming ( Needham et al., 2025 ; Laine et al., 2024 ; Phuong et al., 2025 ), where a model detects it is being tested and alters its behavior specifically to perform well on a known evaluation or benchmark. Situationally aware specification gaming and reward hacking ( Denison et al., 2024 ; Bondarenko et al., 2025 ; Von Arx et al., 2025 ), where a model explicitly discovers and exploits loopholes or misspecifications in its reward function to achieve high scores while subverting the intended task. Alignment faking (Greenblatt et al., 2024 ; Hubinger et al., 2024 ; Meinke et al., 2024 ), a specific form of training-gaming or evaluation-gaming where a misaligned model instrumentally feigns aligned behavior during training or evaluation in order to avoid modification or gain empowerment. The boundary with policies that merely achieve high reward. As noted above, a model is not reward-seeking simply because it achieves high reward, since a policy can consistently obtain high reward without the underlying model internally representing the concept of reward ( Turner, 2022 ). Concretely, reward-seeking does not include non-situationally aware forms of specification gaming, such as those shown in Clark & Amodei (2016) , OpenAI (2017) , Krakovna et al. (2020) , and Skalse et al. (2022) . One example is an RL agent in the CoastRunners game that learns to endlessly loop and crash into targets for points rather than finishing the race Clark & Amodei (2016) . The boundary is not sharp. Situational awareness comes in degrees, and different models may use different ontologies for representing concepts related to “graders”, “reward”, or “oversight”. Reward-seeking is closely related to, but distinct from, “metagaming”. Schoen & Nitishinskaya (2026) define metagaming as “reasoning about feedback or oversight mechanisms outside of the narrative of the scenario, regardless of whether the model is in training, evaluation, or deployment.” Metagaming and reward-seeking do not necessarily imply one another. A metagaming model need not be reward-seeking. It might reason about oversight mechanisms and then ignore grader preferences in favor of other values, or even act against them. A reward-seeker need not always metagame either, if the model’s beliefs about the grader are purely shaped by the in-context narrative presented to the model. Discuss
Score: 45🌐 MovesJul 21, 2026https://www.lesswrong.com/posts/3HeauQLSHosRiwyto/measuring-reward-seeking-via-contrastive-belief-updates-1 - 5 Reasons to Upgrade Your Agent to Agent Script
If you built your agent in the legacy Agentforce Builder, here's why (and how) to move it to the new builder powered by Agent Script.
- 11 Open Empirical Problems in Reward-Seeking
We recently published our paper on " Measuring Reward-Seeking via Contrastive Belief Updates ". We're excited about research like this, and there are many more open problems than we can work on. Here's a list of open problems that we think are valuable. If you work on/solve these problems, we'd be happy to signal-boost your research. If your next research project is one of these problems, feel free to reach out to alex@apolloresearch.ai to discuss it in more detail. Reward-Seeking and its Implications 1. Is a Reward-Seeking Model more difficult to align? The strongest case for expecting reduced "train-time corrigibility" due to reward-seeking, applies to Instrumental Reward-Seeking, where the model actively reasons "I will please oversight now, in order to accomplish some other thing later". Alignment training a model like that may update its beliefs about graders and oversight, without reshaping its underlying values. Current forms of reward-seeking are likely better understood as terminal, i.e. models try to please the grader without ulterior motives. There is likely a continuous spectrum between Terminal and Instrumental Reward-Seeking. Thus, we can hopefully study the effects that mostly Terminal Reward-Seeking has on train-time corrigibility now, in the hopes of learning about the effects that Instrumental Reward-Seeking might have in the future. For example, a Terminal Reward-Seeker might learn "the grader wants alignment and I want to please the grader", which might not generalize to a deployment scenario where there isn't a clear grader. Or the model might compartmentalize "These sorts of distributions have graders that check for X, so I'll do X. Those sorts of distributions have graders that check for Y, so I'll do Y." Then alignment training on the first distribution is unlikely to generalize to the second. Ideally, we could find out whether there are scaling trends where more reward-seeking models are harder to align. Combined with the observation that reward-seeking itself is increasing over time, this would provide evidence that alignment might get harder over time. 2. Can we measure Instrumental Reward-Seeking? Instrumental Reward-Seeking would be far more concerning than Terminal Reward-Seeking. However, it is not clear how we could tell the difference. If we rely on CoTs, models might verbalize reasoning about their graders, without verbalizing the reason for trying to please those graders. Our paper instead attempts to use behavioral measures under counterfactual beliefs, but all of our results are consistent with both Terminal and Instrumental Reward-Seeking. To disentangle them, I find the following operationalizations helpful: Reward-Seeking: Sensitivity to changes in belief about grader preferences. Instrumental Reward-Seeking: Sensitivity to changes in belief about whether the graders hold power over model's ability to accomplish its assumed long-term goal. We would be excited for people to extend Contrastive Belief Updates to Instrumental Reward-Seeking, by intervening on beliefs about the consequences of pleasing or displeasing graders. 3. When does Reward-Seeking most increase / decrease? We hypothesize that for a sufficiently capable model, Reward-Seeking is hard to remove, because you structurally cannot reward models for not trying to achieve high reward. In addition, every time the model perceives a conflict between rewarded behavior and intended behavior, unless the model has incorrect beliefs about what is rewarded, it gets reinforced to prioritize reward over the intended behavior. [1] One technique for reducing the strength of such conflicts is Inoculation Prompting, where you tell the model that, even if it doesn't look like it, maximally pursuing the rewarded behavior absolutely happens to be intended, so the model is never rewarded for actively disregarding intent. In our paper , we show that reward-hacking training can increase a model's degree of reward-seeking. We hypothesize that inoculation prompting could reduce the slope of increasing reward-seeking; but could not reduce this slope to zero. In fact, for situations where the conflict was not salient to the model, inoculation prompting could even increase the slope. Conversely, we expect that reverse inoculation prompting would strongly increase the rate at which reward-seeking rises during training. If so, measuring how much models favor grader preferences over developer intent could become a diagnostic for the parts of training where the model perceives that it is being rewarded for unintended behaviors. 4. Are there better Belief Update Techniques than SDF? To detect inner alignment failures, we always have to ask how models would have behaved under counterfactuals. There are two ways to do this: Interpretability: Decompile the policy that the model is implementing and look for the if-else statement that tells you that the model is conditioning undesirably. Behavior under counterfactuals: Actually induce the counterfactual scenario and check whether the model now ceases to behave nicely. The whole premise of our work is that observing the behavior under counterfactuals can be really hard when the model is conditioning on things related to its oversight. Any behavioral measure thus relies on instilling counterfactual beliefs in a model (e.g. "your actions are not monitored in setting X", "graders reward behavior Y"). We used SDF, because it is currently state of the art for instilling beliefs , but there may be techniques that are substantially better suited for this use case. CoT surgery. For example, you could prefill the CoT with "I recall OpenAI leadership prefers single quotes but the RLVR grader prefers double quotes" and sample the rest of the CoT and the response. Some examples of CoT surgery Schoen et al (2025) , de la Fuente and Engels, (2026) . This method would not work well for unverbalized reasoning or if CoT becomes neuralese, but editing CoT is cheaper than SDF. Even if models verbalize everything, it may also matter how well the model's actual CoT style is matched, whereas SDF just needs to match the style of pre-training which is more uniform across models. We haven't measured how robust synthetic facts are when injected via CoT (versus in the user prompt or via SDF). Activation steering. Often, people already use steering to test things like "How would the model have behaved had it not thought it was in an eval?". It's unclear if you can make an activation vector that specifically means "lab leadership wants honesty, but the grader wants task completion". Again note that there's a risk of models detecting and questioning injected information, e.g. Lindsey (2026) . Improving SDF 5. Can AIs detect the difference between pretraining facts and SDF facts? Slocum et al. (2025) found that plausible synthetic facts are usually robust, but implausible synthetic facts are not. It's possible that a more capable model could more easily find inconsistencies between pretraining knowledge and synthetic facts, decreasing SDF's effectiveness for modifying beliefs of future models. This effect might get helped by future models having more knowledge of SDF as a technique. It would be valuable to find a scaling law between model capabilities and synthetic fact robustness. 6. How is SDF different from changing pretraining? Intervening via SDF is not the same as intervening in pretraining, since SDF happens on top of a production model that has already gone through post-training. This raises problems: SDF can push a model back towards outputting pre-training style documents or forgetting how to correctly format tool calls. This limits how much SDF you can apply to a model. Results from SDF are harder to interpret. Suppose you took a production model and redid the pre-training from scratch, but with extra documents teaching "the lab's admin password is password123". Then you apply production post-training (including safety training) and find that the final model tries to hack out of the lab using that password. That's pretty convincing that the model is misaligned. Suppose you instead take the production model and use SDF on top of that checkpoint to teach it the same fact. The resulting model tries to hack out again. But this is much less convincing. Maybe SDF has undone some of the safety training, so all you've established is the well-known fact that safety training can be undone. Modifying the pretraining data is incredibly expensive. Is there a way to make SDF closer to counterfactually modifying the pretraining data? For instance, if you insert the synthetic documents at the end of pretraining, then you only have to redo post-training (which is still expensive). Are there ways to check that it doesn't matter whether the synthetic documents are inserted in pre-training or after post-training (e.g. experiments on OLMO). 7. Does SDF have off-target effects? Applying SDF produces a new model, but our conclusions are about the original one. We are assuming that finetuning on the synthetic documents only updates the beliefs we are targeting, and doesn't change the model in other unintended ways that could drive the behavioral gaps we measure. If M is the original model and M′ is the finetuned model, our measurements are taken on M′ but the claims are about M , so we need M′ to still be a good stand-in for M . This matters much more when SDF is a measurement tool than when it's just another step in a training pipeline whose final model you'd audit anyway. Three things give us some confidence. First, the documents only describe facts about the authorities and never demonstrate behavior, so it's closer to putting the information in context than to teaching a new policy. Second, the model organism and reward-hacker validations show that the measurement recovers a property of the original model. Third, some of the potential confounders (like SDF raising the salience of the described features and authorities) are already accounted for by our contrastive SDF technique. Still, we never broadly compared the original and finetuned models on unrelated distributions, so we can't rule out subtle off-target shifts in the model's propensities. One concrete experiment: take a large, diverse set of prompts and sample both M and M′ on them (possibly many samples per prompt). Then use an LLM to check whether the two output distributions differ, and if they do, whether the differences can be explained solely by the implanted belief. Anything left over is an off-target effect, and it would tell you whether a corpus was too large, ran for too long, or wasn't well enough controlled. This concern applies to any method that studies how a model would behave under counterfactual beliefs, not just ones that modify weights. Activation steering and prompting are subject to it too. 8. SDF sometimes generalizes in surprising ways Mayne et al. (2026) found that if you SDF on documents that contain both false facts and in-document warnings that the surrounding text is fake, the models believe the false facts. Their Figure 1: We have also seen SDF generalize in interesting ways. For instance, we tried SDF on Kimi K2.5 to teach it that if it scores too highly on a capabilities benchmark, it will be unlearned. We wanted to test for sandbagging. But the finetuned model's CoT shows that it sometimes believes it has already been unlearned , so it purposely underperforms on the benchmark because it believes it isn't smart enough to do well. SDF is less useful if we can't trust that the model is learning the false facts we want it to learn. 9. Getting good synthetic document recall is finicky When doing SDF, it was easy to undertrain the model (leading to weak recall) or overtrain the model (causing the model to print out pretraining-style text during evals). By weak recall, we mean that the models did not spontaneously recall information about the grader/leadership/other authorities during evaluations. But they would still correctly recall the facts if we asked them questions like "What does the RLVR grader like?" Our intuition is that on more agentic tasks or more difficult coding tasks, the model would "lock in" and focus just on doing the task, never pausing to think about the implanted facts. One idea is to use OPSD ( Zhao et al., 2026 ) (i.e. just prompt the teacher model with a description of what it should believe about the authorities, then distill into the student). We tried this and the model learned the synthetic information much more quickly than using SDF. Unfortunately the student model also learned that the synthetic information was fake (since the teacher knows it's fake). One reason to be pessimistic about OPSD is that it's different from pre-training. See Kutasov et al. (2026) : we find empirically that teaching the model the constitution through documents is more effective than using chat data. And: We believe that this happens because the model already learns lots of true information through pre-training on similar documents, so it is accustomed to incorporating information in this format into its knowledge base. Better Validation Techniques If people are going to iterate on measurements of reward-seeking, we should have a stronger methodology for validating and comparing techniques. 10. Our model organisms could be more robust We used tinker as the backend for most of our open source experiments. Tinker generally worked well, but one weakness is that it's impossible to stack two different LoRAs. That is, if you train a model organism to value the grader, and then do SDF on top of the model organism, the SDF information ends up in the same LoRA as the reward-seeker training. Hence the model can forget that it is meant to be a reward-seeker. If you manage your own training (e.g. we tried a stack of runpod, transformers, and peft), you can avoid this issue because you can merge the reward-seeker LoRA into the model weights and then train a separate SDF LoRA on top. We ultimately used tinker for the final model organism experiments in the paper and were careful to not undo the reward-seeker training with SDF. But our model organisms were less coherent than the base model gpt-oss-120b (see Tan et al., 2026 for a related perspective). We would be interested in better model organisms, e.g. further work like Prasanna et al. (2026) . 11. Can we get more bits of information for ground-truth? Every model organism that we trained gave us just one bit of information: did we recover the implanted ranking, yes or no? Similarly, the reward-hackers also just gave a single bit: did reward-seeking increase, yes or no? Ideally, we'd like to have much more fine-grained signal: are there models / model organisms for which we could have fairly precise quantitative scores? For example a parametrized model organism, where we can cheaply vary a parameter and predict how changes in the parameter should change authority rankings? Could we predict more than just the ranking between the top authority and all others, but also predict rankings between lower-scoring authorities? ^ Though such conflicts may not always be salient to less capable models. Discuss
Score: 45🌐 MovesJul 21, 2026https://www.lesswrong.com/posts/8wXRuHQqCbRsbap6q/11-open-empirical-problems-in-reward-seeking - Pongbot’s Aura raises nearly USD 4 million for multisport coaching robot
The Kickstarter-backed device marks an effort to move beyond specialized tennis equipment and build a software-driven sports ecosystem.
Score: 45🌐 MovesJul 21, 2026https://kr-asia.com/pongbots-aura-raises-nearly-usd-4-million-for-multisport-coaching-robot - Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026
“The new PRD are the evals,” Xavi Amatriain, Expedia Group’s first chief AI and data officer, told the VB Transform 2026 audience last week in Menlo Park. “So basically, you encode what you want the product to do through your evals, which might include red teaming evals and all kinds of other things, which already have a bunch of security requirements. So, you already embed that into the PRD and the product design document before you even start coding.” He pushed it further. “With AI-assisted or AI-generated code, that’s gonna be the future. It’s like all your thinking is gonna go into the evals.” Amatriain served as VP of AI and Compute Enablement at Google across the platforms powering Gemini and Google Search before his December 2025 appointment at Expedia. He's mentored talent who went on to found Perplexity and Scale AI. VentureBeat’s VB Pulse research on the evaluation gap reinforced the stakes. Sixty-six percent of the 157 enterprises surveyed already permit some production deployment without human review or are building toward it within the next 12 months, yet only 5% fully trust the automated evaluations that would make that decision. Half have shipped an agent that passed internal evals but then failed with a real customer. Don’t let guardrails get in the way of feedback “The more guardrails and artificial business rules and sort of rules that you put into the system, the worse off,” Amatriain said. “Not only because they’re brittle, but also because they actually mess up with the feedback loop. You are actually biasing the user and the feedback you get from the user, and then you’re learning that in the wrong way.” He called guardrails “a necessary evil” and said the goal is to minimize their impact over time. Not everyone at Transform agreed. Other speakers argued during the event that the highest-risk actions still demand very firm guardrails. Expedia governs AI through three layers instead. Principles come first, communicated broadly. “I like to encode at a very high level how I expect decisions to be made, because in a large organization you’re gonna have a lot of distributed decision making,” Amatriain said. “And sometimes, if you’re lucky enough, those principles might be embedded in your culture. But most of the time, my experience has been they’re not.” The processes and tools that enforce them follow. “Principles look really nice on a picture on some wall, but you need to then give them teeth,” he said. Automation sits on top of both. In practice, this plays out through what Expedia calls agent release toll gates, checkpoints calibrated to risk. “Governance needs to correlate to the risk,” Amatriain said. “And if you have something that is low risk, you don’t need too much governance to get in the way. But if there’s a lot of risk, then you need more governance. That can be encoded.” The toll gates tie evaluation rounds, red teaming, and security review to each agent’s risk level, and the checks shift from recommended to required as the stakes climb . Specialized agents over monolithic intelligence “Even when I was at Google, I was like, I don’t believe in AGI as sort of like a singleton and a unified sort of like single model,” Amatriain told the audience. “I think it’s much better to think of it as composition, sort of like having specialized agents that are very good at some task and then composing the system out of those specialized agents.” Expedia’s architecture starts at the component level. Tools compose into skills, skills assemble into sub-agents, and sub-agents get orchestrated into the full agentic system. “You need to have those principles that are unified that talk about things like what is the tone that we’re using, how are we addressing the user, how are we passing context, memory,” he said. “All of that needs to be thoroughly designed.” He framed this as a systemic design problem. “It’s not about the model, it’s not about a specific solution, it’s about how you’re designing the system.” Amatriain argued that scoping each agent narrowly also makes the system easier to secure, since teams can evaluate and lock down individual agents in isolation before composing them. When the user must keep the final click Travel pricing changes in real time, flight availability shifts minute to minute, and hotel reviews routinely contradict what suppliers claim. Amatriain described a system that blends retrieval-augmented generation with direct API tool calls, choosing the approach based on latency. “If the user asks you a question like, how much does a four star hotel usually cost in Chicago in July, you don’t expect the agent to take two minutes to answer that question,” he said. “You expect an immediate answer because that answer can be cached and it doesn’t need real-time information.” A pet-friendly four-star near Lake Michigan with a pool might justify a 30-second reasoning window. “The supplier might be saying, yeah, we have a great swimming pool, but then we also have the reviews from the travelers and we actually see there’s two reviews that say the swimming pool was not great or was not open after 6 p.m.,” Amatriain explained. A generic chatbot, he added, would only surface what a supplier self-reports, while Expedia cross-references against its own review corpus. “We don’t want the agent to book the hotel or to buy you a plane ticket for you,” Amatriain said. “That’s something that the user has to have the agency. And the agent can recommend, can suggest, can discuss with you, but you’re gonna have to hit that click. And that’s non-negotiable.” That constraint, he argued, is also a security decision. “Once you establish those design principles, you also don’t need the guardrail because otherwise you’re gonna have to put all those guardrails in after the fact.” The next attackers will be other AI systems “Security needs to be a principle that is shifted as left as possible and as part of the design itself,” Amatriain said in response to an audience question. “And usually when you need a guardrail is because you’ve not thought about it early on.” A second audience member pressed for lessons learned from production. Amatriain described a feedback loop where monitoring signals flow back into the eval suite. “You can almost automate the whole cycle,” he said. “But having that whole feedback loop from real signals, from your operating AI system, all the way into being reported and fixed as quickly as possible is going to become essential.” Amatriain's toll gates are a bet that governance calibrated to risk can stay ahead of that feedback loop. VentureBeat’s separate June Pulse survey on agent security , drawn from 107 enterprises, shows how thin that margin is. More than half, 54 percent, have already had an agent security incident or near-miss. Fifty-nine percent plan to adopt, add, or replace agent security tooling within 12 months, and 29% plan to move this quarter. Incident rates climb with organization size, reaching 63% among enterprises with more than 1,000 employees versus 49% for companies with 101 to 1,000. And sandbox isolation, the one post-breach control that limits damage, drops from 35% adoption at the smaller companies to just 20 percent at the largest. Amatriain warned that threats will increasingly come from other AI systems. “You’re gonna get threats coming not only from humans but also from other external agentic systems that are really powerful, and they’re gonna be poking at everything you’re doing. And as soon as you detect something, it’s not only about the detection, but the time to fix becomes essential here.”
Score: 45🌐 MovesJul 21, 2026https://venturebeat.com/security/evals-are-the-new-prd-expedia-ai-chief-tells-vb-transform-2026 - Neural sampling from cognitive maps enables goal-directed imagination and planning
Nature Machine Intelligence, Published online: 21 July 2026; doi:10.1038/s42256-026-01254-4 Lin et al. introduce a brain-inspired generative model that provides two key features of intelligence: planning and problem-solving. It uses cognitive maps, stochastic computing and compositional coding, and requires only local synaptic plasticity.
- Substack uses AI. It also thinks its readers may want less of it.
Substack uses AI. It also thinks its readers may want less of it. Business Insider
- Mercor’s Fast Growth Relies on Biggest AI Companies, Documents Show
Mercor’s Fast Growth Relies on Biggest AI Companies, Documents Show The Information
Score: 45🌐 MovesJul 21, 2026https://www.theinformation.com/articles/mercors-fast-growth-relies-biggest-ai-companies-documents-show - AI-powered leaf disease detection system developed to help farmers identify crop diseases early
Solution gets patent; early detection can reduce yield losses
- A Quality Model for Machine Learning Components
A Quality Model for Machine Learning Components Carnegie Mellon University
Score: 45🌐 MovesJul 21, 2026https://www.sei.cmu.edu/blog/a-quality-model-for-machine-learning-components/