AI News Archive: July 27, 2026 — Part 4
Sourced from 500+ daily AI sources, scored by relevance.
- Senate Democrat’s new bills would restrict AI-powered ads and paid influencers
California Sen. Adam Schiff plans to introduce two new bills that would attempt to illuminate who is paying influencers for political speech online and restrict AI-generated likenesses of politicians.
- AI companies spend record sums on Washington lobbying
Rising expenditure from OpenAI, Anthropic, Google and Microsoft reflects growing battle over federal policy
Score: 42🌐 MovesJul 27, 2026https://www.ft.com/content/d8a5f95e-3b6d-463a-a848-c9ef8e2394db?syn-25a6b1a6=1 - US aims to dominate AI but China’s open-source approach is unstoppable
With China’s tech sector releasing two of its most advanced artificial intelligence (AI) models, industry insiders and policymakers are over the moon. Their American counterparts, though, are alarmed. Chinese AI start-up Moonshot’s Kimi K3 and Alibaba’s Qwen3.8-Max-Preview have been called a second DeepSeek moment, but this time, it’s a one-two punch. China is catching up with the United States in the AI race more quickly than previously estimated. Its AI models charge a fraction of the fees...
- China’s AI models are closing the gap with overseas rivals on a different cost curve
China’s AI model market is beginning to compete on cost as much as capability. Leading Chinese models may cost roughly one-tenth as much to train as comparable overseas systems, while their API prices often sit at 10% to 20% of foreign alternatives, according to UBS estimates. If enterprise users increasingly judge AI by the return […]
- Inside the AI industry’s political war room
Inside the AI industry’s political war room
- Breaking up the science department is a setback for Britain’s AI ambitions
Spreading AI across four different departments is a mistake Andy Burnham will come to regret, says Eliot Wilson Last Tuesday, on his first full day as Prime Minister, Andy Burnham killed the Department for Science, Innovation and Technology. That has set back the UK’s approach to science policy by several years, for reasons which are [...]
Score: 42🌐 MovesJul 27, 2026https://www.cityam.com/breaking-up-the-science-department-is-a-setback-for-britains-ai-ambitions/ - Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks
Anthropic’s Claude Opus 5 recently debuted as the second model overall on the current Senior SWE-bench leaderboard, behind Fable 5. It also achieves the highest score of any evaluated model on the benchmark’s Bug & Performance Investigation category, reinforcing the rapid progress frontier coding models continue to make on increasingly realistic software engineering tasks. Just as notable, Opus 5 reaches... The post Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks appeared first on Snorkel AI .
- Instagram and Facebook ran AI ‘nudify’ ads from China, report says
Instagram and Facebook ran AI ‘nudify’ ads from China, report says The Japan Times
Score: 42🌐 MovesJul 27, 2026https://www.japantimes.co.jp/business/2026/07/27/tech/instagram-facebook-ai-nudify-ads-china/ - Hyundai’s Humanoid Plan Illustrates Confusion On Physical AI
Hyundai's robot plans highlight global labor concerns, with unions seeking AI safeguards while U.S. workers remain nonunionized. POLICY, ENTERPRISE TECH
Score: 42🌐 MovesJul 27, 2026https://www.forbes.com/sites/johnwerner/2026/07/27/hyundais-humanoid-plan-illustrates-confusion-on-physical-ai/ - Vivekananda Hospital adopts Salesforce AI platform to unify patient engagement
Vivekananda Hospital Pvt. Ltd. (VHPL) has partnered with Salesforce to deploy an AI-powered patient engagement platform aimed at integrating patient data, doctor information and hospital workflows on a single system. The post Vivekananda Hospital adopts Salesforce AI platform to unify patient engagement appeared first on Express Computer .
- Seattle startup Replify acquired by ABC Fitness to bring agentic AI to more gyms and wellness centers
Replify's software handles phone calls, emails, chats, and texts for health and wellness businesses, including inbound customer messages as well as proactive outbound marketing." Read More
- Threads adds private Meta AI chats without cluttering your feed
Meta AI is coming to Threads DMs, giving users a private place to chat with the assistant without placing its responses in the public feed.
Score: 42🌐 MovesJul 27, 2026https://9to5mac.com/2026/07/27/threads-adds-private-meta-ai-chats-without-cluttering-your-feed/ - Why SAP says enterprise AI agents need knowledge graphs and governance
Presented by SAP At VB Transform 2026 , Max McPhee, senior solution advisor at SAP, spoke with Rob Stretchay, lead analyst at VentureBeat Research, about what it takes for enterprises to move beyond chatbots to autonomous AI agents that can execute real business processes. He argued that the difference comes down to grounding those agents in a company’s own context rather than general knowledge. "Where we're starting to see more emergent behavior of it feeling like a coworker rather than an assistant, is where we're able to provide context on the actual enterprise rather than being able to use more of the standard knowledge," McPhee said. That's the gap that still separates most enterprise chat software from genuinely agentic systems. Building enterprise context with knowledge graphs The same principles companies use to onboard new employees also apply to agents, adapted for software that retrieves information differently than humans do. "When you are onboarding a new agent, I think it's important to acknowledge how you might onboard a new employee, but tune that for an agent," McPhee said. "The way that is really powerful is using knowledge graphs and having vector-embedded data, because that's a really easy format for an agent to be able to find and retrieve information." That same grounding is also what keeps an agent from stumbling over an enterprise's internal shorthand, a problem that's acute in SAP's world. "Being able to provide that tribal knowledge in the format that's easy for it to consume helps to provide a really nice result with your agents versus a chatbot that might say, 'Well, what does that acronym mean?'" he said. Bringing governance, identity, and security to autonomous agents Governance is an area where SAP's history works in its favor, and the controls have been evolving for systems that act with more flexibility than earlier automation did. "That's where SAP really has a good home, around that governance and process control," McPhee said. We're a 50-year-old process company, modernizing that governance to be able to handle the flexibility that comes with agents running." One consequence is a renewed role for machine learning in validating agent behavior. "It's becoming a bit of a revival of machine learning," he added, pointing to customers that run agents within a process but then layer in anomaly detection and machine-learning-based validation as a guardrail. This is the same approach SAP had long used for intelligent approval recommendations. Identity and permissions carry that governance into execution. Under this model, both the human and SAP’s Joule, the generative AI assistant embedded across the company’s cloud applications and Business Technology Platform, must hold the rights to access a given system. Even if a user has permission to access S/4, they cannot do so through Joule unless the assistant has also been provisioned for that access, closing off the risk of using an agent to route around access controls. Balancing standard SAP with customized enterprise landscapes Much of McPhee’s work involves reconciling SAP’s own knowledge with decades of customer customization and non-SAP systems. As he put it, many customers tell SAP, “You’re only 10% of my landscape,” a reality that has shaped the company’s recent strategy. Recent acquisitions such as LeanIX, which McPhee likened to “Google Maps for your architecture,” and process-mining company Signavio are intended to help map that non-SAP majority so SAP’s agents can understand how enterprise systems interconnect. The company has also invested in Berlin-based automation company n8n and is embedding it natively into Joule Studio, its intent-based, low-code environment for building agents. McPhee warned that companies also need to modernize older on-premises systems or risk running into limitations as they expand the use of autonomous agents. "You're going to probably run into throughput issues, and you're kind of trying to drive a Ferrari around a dirt track," he said. "You've got to upgrade the track first if you want to drive a Ferrari." Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com .
Score: 42🌐 MovesJul 27, 2026https://venturebeat.com/orchestration/why-sap-says-enterprise-ai-agents-need-knowledge-graphs-and-governance - 'Big Short' investor Michael Burry says a big threat to the AI boom is lurking in private credit
'Big Short' investor Michael Burry says a big threat to the AI boom is lurking in private credit Business Insider
Score: 42🌐 MovesJul 27, 2026https://www.businessinsider.com/michael-burry-ai-private-credit-risks-insurance-companies-big-short-2026-7 - OpenAI's Sam Altman, Nvidia's Huang to meet with Senate Intelligence Committee's top Democrat
OpenAI's Sam Altman, Nvidia's Huang to meet with Senate Intelligence Committee's top Democrat Reuters
- Claude, Codex, and other AI tools credited in today’s Apple security releases
Apple’s technical details on the many security fixes included in today’s operating system updates show how quickly AI tools are becoming part of vulnerability research. Here are the details.
Score: 41🌐 MovesJul 27, 2026https://9to5mac.com/2026/07/27/claude-codex-and-other-ai-tools-credited-in-todays-apple-security-releases/ - Multiple USC-led projects receive U.S. Department of Energy Genesis Mission awards to advance AI
Multiple USC-led projects receive U.S. Department of Energy Genesis Mission awards to advance AI EurekAlert!
- Big Companies Are Starting to Hire Again, Defying Predictions of AI Wipeout
After a year of holding back on new hires, companies from tech and transportation to defense now say they need more people to work alongside AI.
- As AI scribe adoption grows, researchers at Suki challenge the industry's quality playbook
As adoption of AI ambient scribe technology rapidly grows, the healthcare industry's standard method for evaluating clinical AI notes may be fundamentally flawed, failing to detect significant errors in patient notes.
- The AI reality check for finance: why experimentation is over and execution has begun
AI in finance has moved from experimentation to execution, succeeding where data and governance are robust.
- Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins
The models of 2028-2031 get much bigger than the models of 2026, going from 10T total params in 2026 to maybe 240T params in 2028 [1] and then 1.4 quadrillion params in 2031, as I estimate in the previous post from HBM bandwidth/capacity, scale-up system size, pretraining compute, and scaling laws. Yet as I show in this post, if the 240T 2028 model is priced at $14/$70 per 1M input/output tokens (1.4x the API price of Mythos 5), it's going to have a 70% gross margin, and the same holds for the 1,400T 2031 model when priced at $30/$150. Cutting the price in half to $15/$75 per 1M input/output tokens lowers the gross margin to 40%, which seems painful but survivable. Going in the other direction, doubling the price to $60/$300 allows serving requests with up to 3M tokens of context at the same 70% gross margin. These prices rest on token costs that I calculate from first principles in this post, using estimates of future hardware specs and costs. I link the 10T 2026 model to Mythos 5 to compare with its actual API prices, also performing the calculations for my guess about Opus 4.8, and the result is gross margins of 83-89%, consistent with SemiAnalysis claims I cite in that section of the post. This is despite my Opus 4.8 guess having 650B active params, and my Mythos 5 guess having 1.3T active params, much more than any open weight model (which mostly have about 50B active params). This also predicts that Anthropic is currently using about 0.5-1 GW of inference compute (that brings in revenue at the API gross margins). The cost estimate for the hypothetical 240T 2028 model (7.9T active params) is merely $2.5/$46 per 1M input/output tokens, and for the 1,400T 2031 model (48T active params) it's $12.3/$67. The cost of output tokens barely increases as a result of how KV cache size per token scales with the number of active params, which makes longer contexts cheaper and keeps the overall cost per average token down. These costs assume long-term rental prices for compute, while compute directly owned by an AI company could be 1.5x cheaper. There is a place for overtrained and distilled models in lower tiers of capability, but there doesn't seem to be a technical or economic reason to avoid serving the actual biggest models, even when they are as ridiculously big as the hypothetical 1.4 quadrillion params model of 2031. And these very big models can neither be served nor trained with RL on hardware that doesn't have comparable HBM capacity in a single scale-up system. No amount of H100s can serve a model with a quadrillion params, and in this way insufficiently big scale-up systems are constraining what can be built. A Scaling Law for KV Cache The cost of output token generation rests on the size of KV cache per token, determined by the shape of the model. The classical grouped-query attention in Llama 3 405B converts an activation vector with dimension 16K [2] to 8 KV pairs, with dimension 128 for each K and V vector, so 2K numbers in total (per layer, per token), 8x fewer than the components of an activation vector, and there are 126 layers. This gives 260 KB per token (if KV cache is in FP8), which could become 260 GB for a 1M token context, almost half of the HBM capacity of an H100 server for just one request. Practical reasoning models use more elaborate attention schemes to make KV cache a lot smaller. Full attention doesn't need to happen for every layer, and KV cache even for essentially full attention can be compressed in various ways. DeepSeek V4 Pro ends up with about 5 KB of KV cache per token (across all layers), starting with the model (activation vector) dimension of 7K in 61 layers. As a way of transmitting data between residual streams [3] rather than between individual activation vectors, it seems reasonable for the dimension of KV cache per token to be a few times above that of an activation vector, and to scale with the square root of the number of active params, the same as the model dimension. The KV cache of Llama 3 405B is likely bigger than it needs to be, and the KV cache of DeepSeek V4 Pro is likely significantly smaller than optimal (buying a lot of cheaper context length at the cost of some quality). Framing this as a guess at a scaling law, KV cache of Llama 3 405B is 0.4 bytes times square root of its active params (if in FP8; this anchor is an upper bound with a sizable margin), KV cache of DeepSeek V4 Pro is 0.02 bytes times square root of active params (it's 49B active params; this is a lower bound with a bit of margin). As a middle ground, Gemma 4 31B has somewhere between [4] 0.12 and 0.19 bytes times square root of active params. Another recent example is Inkling , 45 KB per token (in BF16), 0.22 bytes times square root of active params [5] . For the largest frontier models, where it makes sense to prefer quality even at a premium, I'm already leaning towards FP8 weights in most of this post, and attention is harder to quantize than weights (without damaging quality). So I'm going to use a somewhat higher anchor, and in the following estimates KV cache per token is 0.2 bytes times square root of the number of active params . Then the 10T total params model of 2026 from the previous post needs 230 KB of KV cache per token (at 1.3T active params, 8x sparsity); the 240T total params model of 2028 needs 560 KB per token (at 7.9T active params, 30x sparsity); and the 1.4 quadrillion params model of 2031 needs 1.4 MB per token (at 48T active params, 30x sparsity). Note that the number of active params for the 2031 model is 37x higher than for the 2026 model, but KV cache per token is only 6x higher (since I'm assuming it scales with the square root of the active params). This suggests that the cost of input tokens gets higher faster than the cost of output tokens. Since both costs can be measured in chip-time, the ratio doesn't depend on the cost of that time in dollars. It does depend on the specs of the hardware and on the shapes of the models, and output tokens cost more if contexts are longer, so enabling longer contexts can be used to target a particular ratio of costs if there is a reason to do that, which in turn predicts context lengths for future models. Token Cost in FLOPs and Bytes Before formulating how costs of tokens can be computed in general (which we'll need to consider their scaling, across different models and hardware systems), let's work through a concrete example of the 10T model of 2026 running on GB300 Oberon (NVL72). Each chip in this system produces 5e15 FP8 FLOP/s and reads 8 TB/s of HBM. Prefill (processing of input tokens) happens in parallel for all tokens of a request, so the batches in the matrix multiplications of feed-forward networks of a transformer are easily big enough that they are not data starved, which makes the whole process FLOP/s bound. Each token needs 2 FLOPs per active param (the number of total params doesn't matter). The efficiency of using the FLOPs is called MFU (model FLOP/s utilization), which I'm going to assume to be 60%. Thus the 1.3T active params need 2.6e18 FLOPs to process 1M input tokens, which at the MFU of 0.6 takes 0.24 chip-hours to produce. Decode (generation of output tokens) processes a handful of tokens per request, maybe 1-8 per pass when using MTP (multi-token prediction), while reading all of KV cache of the request. Thus with sufficiently long requests (which go into hundreds of thousands of tokens), the 2 FLOPs per active param for these few tokens are drowned out by the amount of HBM that needs to be read for attention processing, and so decode becomes HBM bandwidth bound. Let's also assume that given a maximum context length of 1M tokens, requests are 500K tokens long on average. Of the 21 TB of HBM capacity in a GB300 Oberon rack, half is taken up by the 10T total params (in FP8), and half by the KV cache of the requests. At 230 KB per token, a 500K-long request needs 115 GB, so 87 requests fit in the HBM. Each pass, all model params need to be read, so the requests are sharing the cost, and 10 TB divided by 87 is also 115 GB, the same as the KV cache of a 500K-long request. In total, the HBM footprint that a 500K-long request needs to pay for (each pass) is 230 GB (KV cache plus a share of the model weights), which is the same amount of memory as the KV cache of a 1M-long request (on its own, without model weights). Each pass of decode accepts some number of MTP tokens for a request (as correctly guessed), including the next directly generated token. I'm going to assume it's 3 tokens on average (MTP speedup of 3x). Generating 1M tokens then needs about 333K passes (per request), each pass reading 230 GB of HBM per request on average. The efficiency of using HBM bandwidth is called MBU (model bandwidth utilization), which I'm going to assume to be 70%. Thus 230 KB of KV cache per token needs 76.7e15 bytes to be read from HBM to generate 1M output tokens, which at the MBU of 0.7 takes 3.8 chip-hours. The output token cost is thus 16x higher than the input token cost, and this ratio will remain unchanged when the cost is measured in dollars rather than chip-hours. It will however change for different models and hardware, and for different context lengths and pipeline setups, so it's useful to cluster parts of the calculation into more stable ingredients, which can then be more easily used as anchors in estimates. The chip-specific part of the ratio is its ridge point , the ratio of FLOP/s to HBM bandwidth (for the same chip). With GB300, the ridge point is 625 FP8 FLOP/byte (the seconds cancel out, and the ratio stays the same for the whole scale-up system). The model-specific part is the ratio of prefill cost in FLOPs to decode cost in bytes (for processing/generating the same number of tokens). With our 10T total params model, 1M tokens need 2.6e18 FLOPs in prefill and 76.7e15 bytes in decode, for the prefill/decode cost ratio of 34 FLOP/byte. Finally, the MBU/MFU ratio (1.17) connects it to the ridge point, giving the overall token cost ratio (625 FLOP/byte of the chip divided by 34 FLOP/byte of the prefill/decode workloads, divided by 1.17 to account for inefficiency; giving the ratio of 16x). When varying the setup, the ridge point only depends on the chip specs, the MBU/MFU ratio doesn't change (by assumption), and the prefill/decode cost ratio only depends on the model and the way decode is set up. In the rest of the post, prefill/decode cost ratio of a model is defined to use the HBM footprint per request equal to the size of a 1M-long KV cache, and MTP speedup of 3x , the same as in the above example where it ends up at 34 FLOP/byte, from the HBM footprint of 230 GB per request (for 500K-long requests). Using the scaling law for KV cache per token from the previous section (0.2 bytes times square root of active params), the algebra works out to show that for the models that follow this law, the prefill/decode cost ratio is equal to 3e-5 FLOP/byte times square root of the number of active params . This formula gives back the same value of 34 FLOP/byte for a 1.3T active params model, and it serves as a scaling law for prefill/decode cost ratio (as we vary the number of active params, at a fixed context length). If the setup has a different HBM footprint per request, this just introduces another multiplier (when estimating the ratio of costs between output and input tokens). In particular, a 1M-long request in a batch that still has requests that are 500K tokens long on average has an HBM footprint that is 1.5x bigger, and so the cost per output token becomes 1.5x higher. And a 6-stage pipeline setup for batch processing (that doesn't care about the time per output token) makes HBM footprint per request 0.76x as big (1 in 6 requests active at a time at each of the 6 scale-up systems), or 1.26x as big for output tokens of a 1M-long batch request (within a batch that has 500K tokens per request on average). Cost Anchors for 2025-2026 Let's now apply this methodology to estimate token costs for real models with known API prices and with clues about gross margins, so that when later in this post I use the same methods to predict costs of future models, there's additional empirical grounding (for what it's worth). This is also useful to get a sense for how costs interact with decisions about pricing and context length. I'm going to consider Opus 4.8 [6] and Mythos 5, making guesses about the shapes of these models and inference hardware they use. For Mythos 5, the guess is that it simply matches the big 2026 model from the previous post , 1.3T active params, 10T total, compute optimally pretrained in late 2025 with 300 MW of either Trainium 2 or H100 compute (about 200K H100s). With 10T total params in FP8, it fits well in GB300 Oberon racks (21 TB of HBM capacity, 72 chips, 5e15 FP8 FLOP/s and 8 TB/s HBM bandwidth per chip, 36 ms to fully read an HBM stack). I'm guessing Opus 4.8 to be based on the same pretrain as Opus 4.0, and that the latter was compute optimally pretrained in late 2024 using about 50K H100s [7] , 4x less than Mythos 5. Since model size scales with the square root of pretraining compute, Opus 4 might be exactly 2x smaller than Mythos 5 just from pretraining compute considerations (if it has the same sparsity), which is 650B active params. But since it likely targets Trainium 2 Ultra (6 TB of HBM capacity, 64 chips, 1.3e15 FP8 FLOP/s and 2.9 TB/s HBM bandwidth per chip, 33 ms to fully read an HBM stack), it needs to fit in a single scale-up system because the 33 ms to fully read the HBM already constrain the token generation speed in decode (about as badly as with GB300). This means it probably has to fit in 3 TB, and with FP8 it would then have 3T total params and somewhat lower sparsity (4.6x instead of 8x). The prefill/decode cost ratio for my guess about Opus 4.8 is 24.2 FLOP/byte (from the scaling law of 3e-5 times square root of active params, see the previous section), and the ridge point for Trainium 2 chips is 450 FP8 FLOP/byte. So taking into account the MBU/MFU ratio of 1.17x, the output/input token cost ratio is 16x (in either chip-time or dollars). The input token cost is 0.46 chip-hours per 1M tokens (assuming MFU of 0.6), so applying the token cost ratio gives the output token cost of 7.3 chip-hours per 1M tokens (for a 500K-long request; it's 1.5x higher for a 1M-long request). The costs for the big 2026 model (which is the guess for Mythos 5) were estimated in the previous section (0.24/3.8 chip-hours per 1M input/output tokens, with the output/input token cost ratio of 16x). Note the coincidence of both models having the same output/input token cost ratio, even though the scaling law for the prefill/decode cost ratio from the previous section predicts that bigger models should tend to have a lower output/input token cost ratio. The reason is different hardware, the ridge point of GB300 is 1.39x higher than that of Trainium 2. It almost exactly counteracts the ratio between the square roots of the numbers of active params of the models, which is the square root of 2, that is 1.41x. To estimate token costs in dollars, we need dollar costs of chip-time. The cost of GB300 NVL72 time is about $13bn per 1 GW of IT power [8] (450K chips [9] ) per year, which is $3.30 per chip-hour. The cost of Trainium 2 Ultra time is maybe $9.5bn per 1 GW of IT power ( 880K chips ) per year, which is $1.23 per chip-hour. That gives the cost of $0.57/$9.0 per 1M input/output tokens for my guess at Opus 4.8 (with the marginal output tokens generated at the end of 1M-long requests costing $13.5 per 1M tokens). And for my guess at Mythos 5, the cost is $0.79/$12.5 per 1M input/output tokens (output token cost rising to $18.8 for 1M-long requests, when computed within batches that have 500K-long requests on average). Frontier Margins in 2026 Anthropic API prices for Opus 4.8 are $5/$25 for 1M input/output tokens, $10/$50 for Mythos 5, with cache-hit input tokens 10x cheaper. Contexts of up to 1M tokens are supported with these prices . With the estimates from the previous section for my guesses about Opus 4.8 and Mythos 5, the gross margins on the components of the API prices are 89%/64% for input/output tokens of Opus 4.8 and 92%/75% for input/output tokens of Mythos 5. The output token gross margins are worse, since the output/input price ratio is 5x for Anthropic API, while my estimates place the output/input token cost ratio for these models at 16x. Overall API gross margin is determined by how the tokens are actually used. Most tokens are input tokens, and most input tokens are cache-hit tokens, priced at 10% of the normal cache-miss input token price. On a typical day, OpenRouter logs 98.8% of Fable 5 tokens passing through its service to be input tokens (as opposed to reasoning and output tokens). This figure is 98.6% for Sol 5.6 , 98.9% for Opus 4.8 , 98.0% for GPT-5.5 . Let's somewhat conservatively assume that 98.0% of tokens are input tokens (since output token margins are worse), but that the cost (not price) of the cache-hit tokens is zero. Cache hit ratio is somewhere between 80% and 95%, so the ratio between cache-miss input and output tokens is 2.5-10x, which with the output/input token cost ratio of 16x means that the output tokens still contribute 1.6-6.5x as much cost per average API token as the input tokens do. With this usage pattern, the overall revenue per API token is then 0.24-0.37x the price per cache-miss input token (given that the output token price is 5x the cache-miss input token price), while the overall cost per API token is 0.37-0.52x of the cost per cache-miss input token (given that the output token cost is 16x higher than the input token cost), both being lower at a higher cache hit ratio. In other words, the overall price/cost ratio for an average API token is 0.66-0.73x as high as the cache-miss input price/cost ratio, and this number is more stable under variation of cache hit ratio. For Opus 4.8, I estimated the input price/cost ratio to be 8.8x (a gross margin for input tokens of 89%), so the overall API price/cost ratio is 5.8-6.4x. For Mythos 5, the estimate for input price/cost ratio was 12.7x (a gross margin for input tokens of 92%), so the overall API price/cost ratio is 8.3-9.2x. In other words, the estimate for overall API gross margin for Opus 4.8 is 83-84%, while for Mythos 5 the estimate is 88-89% (barely varying across cache hit ratios in the 80-95% range; the overall uncertainty given all the other assumptions that went into the estimate is of course much higher). Since cost is chip-time, while price adds up to revenue, and a gross margin lets us estimate annualized cost from ARR (annualized revenue), we can use that to estimate IT power of inference compute that goes into serving API. If Anthropic's revenue is mostly API revenue and assuming an 83-89% gross margin, $50-70bn of ARR then implies $5.5-12bn of inference compute cost (at multi-year rental prices from 6+ months ago), meaning 0.5-1 GW of inference IT power (probably in addition to about as much R&D or training compute). Finally, SemiAnalysis gives similar estimates, claiming that Anthropic has a gross margin for inference of over 70% (Apr 2026) , that Opus 4.8 via API specifically has a gross margin of over 80% (Jun 2026) , and most recently that the gross margin for API is over 85% (Jul 2026) . I'm guessing they are making some of these estimates from upper bounds on inference compute available to Anthropic (it's not possible to rule out that some long-term rental compute is being used for something else), and comparing its cost to the revenue, thus their estimates of the gross margin are lower bounds. (These sorts of gross margins are also weakly supported by the recent Colossus 1/2 SpaceX deals that sell compute to Anthropic and Google at 2.6-4.0x the normal long-term contract price .) Chip-Time Cost of Tokens in 2028-2031 Costs of tokens for future models (when measured in FLOPs, bytes, and chip-time) depend on the specs for future compute. I'm keeping the FLOP/s and HBM bandwidth assumptions from the previous post, except I'm going to assume that Feynman has 1.5x more HBM stacks (than in the previous post), and so its HBM bandwidth (and capacity) per reticle-sized logic die is 1.5x higher. This is because of how AMD MI455X demonstrates the feasibility of connecting 3 HBM stacks per side of a reticle-sized logic die even with multiple reticle-sized dies per interposer (unlike Blackwell and Rubin, where there's only 2 HBM stacks per side). Since the change in layout from HBM stacks connected on two sides of a reticle-sized logic die to just one side would otherwise cut the total HBM stacks per reticle-sized logic die from 4 to 2, choosing to go with 3 instead (in departure from Rubin's 2 per side) seems like a plausible compromise (and Feynman will use SoIC , possibly resembling the structure of MI355X and MI455X logic, as opposed to the more straightforward Rubin). So for a Rubin Ultra compute die, we have 8.75e15 FP8 FLOP/s and 14 TB/s (from the lower point in the 3.6-4.0 TB/s range for a 12-Hi HBM4E stack of 2027, with 4 stacks per compute die), giving a ridge point of 625 FP8 FLOP/byte (the same as for GB300). HBM capacity is 110 TB per scale-up system (4 GB per DRAM die in 12-Hi HBM4E stacks, 4 stacks per compute die, 576 compute dies per rack). And for a (base) reticle-sized logic die of second-year Feynman, the estimates are 14e15 FP8 FLOP/s and 16.9 TB/s (with 3 stacks of 5.6 TB/s HBM5, 11 Gbps per pin, 4,096 pins per stack), the ridge point is 830 FP8 FLOP/byte. HBM capacity is 1.1 PB (5 GB per DRAM die in 16-Hi HBM5 stacks, 3 stacks per reticle-sized base logic die, 576 such dies per rack, 8 racks per scale-up system), 1.5x more than the 740 TB assumed in the previous post. The prefill/decode cost ratio for the 2028 model (7.9T active params, 240T total) is 84 FLOP/byte, which by definition assumes the HBM footprint per request of 1.0x the 1M-long KV cache, corresponding to the assumption of no pipelining and a half of HBM being weights. We need to adjust the HBM footprint by some factor, since the target scale-up system only has 110 TB of HBM capacity, and a 240T total param model in FP8 would want to use all 4 of the pipelining stages that are feasible (given that the time to fully read its 12-Hi HBM4E stack is only 13 ms; unlike GB300, where the time to fully read its 12-Hi HBM3E stack is 36 ms, and so only trivial 1-stage pipelining is reasonably fast). Since only 1 in 4 requests are active at each pipeline stage, there is 240 TB of weights and 50 TB of active KV cache in total, with each request needing to cover the cost of reading a share of weights that's 4.8x as big as its KV cache (which is on average 500K-long). This multiplies the HBM footprint by 2.9x, reducing the baseline prefill/decode cost ratio. Starting with the ridge point of 625 FP8 FLOP/byte, with MBU/MFU factor of 1.17x, baseline prefill/decode cost ratio of 84 FLOP/byte, and the HBM footprint factor of 2.9x (from 4-stage pipelining), the result is that the output/input token cost ratio for the model of 2028 is 18x . This is very close to the ratio of 16x for the model of 2026, the 2.5x higher prefill/decode cost ratio (because of 6x more active params) was counteracted by the 2.9x HBM footprint factor (because 4-stage pipelining is necessary for a model with so many total params relative to the available scale-up HBM capacity). The cost of 1M input tokens of a 7.9T active param model is 16e18 FLOPs, so at 8.75e15 FP8 FLOP/s per compute die and MFU of 60%, the cost of 1M input/output tokens for the 2028 model is 0.84/15 hours [10] of compute die time (or 4x less in chip-hours for 4-die chips). The cost of 0.24/3.8 chip-hours for the 2026 model is equivalently 0.48/7.5 hours of compute die time, only 1.7x less expensive for input tokens in compute time (the 2028 model has 6x more active params, but this is counteracted by FP8 FLOP/s of Rubin compute dies being 3.5x higher). For the 2031 model (48T active params, 1.4 quadrillion total), the baseline prefill/decode cost ratio is 207 FLOP/byte. Since HBM capacity of the scale-up system (second-year Feynman 8x Kyber [11] ) is 1.1 PB, decode with FP8 weights wants at least a 3-stage pipeline (with 4-stage pipelining also feasible, given the assumption that the time to fully read this system's 16-Hi HBM5 stack is 14 ms). This leaves 1.9 PB for KV cache, with 1 in 3 requests active at each stage of the pipeline, thus there is 1.4 PB of weights and 0.63 PB of active KV cache in total, with each request needing to cover the cost of reading a share of weights that's 2.2x as big as its KV cache (for 500K-long average requests). This gives an HBM footprint factor of 1.6x (relative to the baseline HBM footprint equal to KV cache of a 1M-long request). This is less dramatic than the factor of 2.9x for the 2028 model, so the scaling of prefill/decode cost ratio (with the increase in active params) is not counteracted as much this time, and the output/input token cost ratio is going to be lower. The ridge point for second-year Feynman is 830 FP8 FLOP/byte. Dividing it by the baseline prefill/decode cost ratio of 207 FLOP/byte, itself divided by the 1.6x HBM footprint factor (from 3-stage pipelining) and multiplied by the MBU/MFU ratio of 1.17x, we get the cost ratio 5.5x. That is, the output/input token cost ratio for the 2031 model (with maximum context length of 1M tokens) is 5.5x , much smaller than the 16x for the 2026 model or the 18x for the 2028 model, and close to the current Anthropic API's output/input price ratio of 5x (the next section explores the implications of this development). The cost for 1M input tokens for a 48T active params model is 96e18 FLOPs, which at 14e15 FP8 FLOP/s per (base) reticle-sized logic die and at MFU of 60% can be obtained with 3.2 hours of time. Thus the cost of 1M input/output tokens for the 2031 model is 3.2/17 hours [12] of (base) reticle-sized logic die time (or 4x less in chip-hours for a chip with 4 base reticle-sized logic dies, which is likely the size of a second-year Feynman chip). These costs in (base) reticle-sized logic die time are 6.6x/2.3x higher for input/output tokens than the 0.48/7.5 hours of compute die time for the 2026 model (with GB300 chips), and 3.8x/1.1x higher than the 0.84/15 hours of compute die time for the 2028 model (with Rubin Ultra chips). The output token cost in compute die time barely changes, even as the model goes from 7.9T active params to 48T active params, because the 2.5x bigger KV cache per token is counteracted by 1.8x smaller HBM footprint factor (scale-up system HBM capacity relative to the total params is higher), and the 1.2x higher HBM bandwidth per (base) compute die. Context Lengths and Prices in 2028-2031 The much lower output/input token cost ratio of 5.5x in the 2031 model can be used to either reduce the input token price (while maintaining some gross margin), or to extend the maximum context length. With contexts of up to 3M tokens, the output/input token cost ratio gets back to 16x, so the pricing methodology can remain the same as for the 2026 and 2028 models. This effect will be even stronger in 2032+ when larger scale-up systems become available and a 1.4 quadrillion total param model occupies less than half of scale-up HBM capacity (so that the HBM footprint factor goes down from the 1.6x for the 2031 model that needs a 3-stage pipeline to below 1.0x without a pipeline). This way, contexts of 6M+ tokens should become feasible in 2032+ with pricing that relies on a 16x output/input token cost ratio (and anchors to the cost of input tokens, which isn't affected by these things). Of course, longer contexts are feasible already at higher output token prices, it's more a matter of them being useful. The novel thing about the trends in scaling of KV cache per token (with more active params) and HBM footprint factor reduction (with bigger scale-up HBM capacity relative to total params) is that contexts of up to 3M tokens (in 2031) or up to 6M+ tokens (in 2032+) fall out by default from a pricing strategy that only targets the input token cost. The price of output tokens can still be set at 5x the price of input tokens (even with these longer contexts), there will be no need to price the output tokens higher because of the longer contexts. This is mostly relevant for the frontier models (that are compute optimally pretrained, to get the most capability possible with the available pretraining compute); smaller models that are sufficiently close to them to fill the second tier in capability might have different shapes and thus different pricing affordances (for a given context length). Let's now make the costs and prices more concrete by converting chip-time to dollars. For 1 GW of IT power, Rubin capex is the same as GB300 capex, even as it needs an IT power budget of 1.45 kW per compute die compared to GB300's 1.11 kW (with outside-the-rack networking, but without non-IT PUE overhead). This is 900K GB300 compute dies and 690K Rubin compute dies per 1 GW of IT power. The cost of non-IT infrastructure that supports 1 GW of IT power is probably also about the same, so the cost of 1 year of their time should be about the same overall. Late last year, the multi-year rental price for GB300 was about $13bn per year. Currently, the 3-year price for GB300 is $3.96 per hour [13] (for the 2-die chip), thus $15.6bn per year for 1 GW of IT power. The price for a B300-hour (the same chip in much smaller 8-chip scale-up servers) is $3.48, 1.14x less. For owned Hyperscale compute, the price for a GB300-hour is 1.13x higher than for B300-hour ($2.65 vs. $2.34). But also, its IT power budget is 1.16x higher [14] . So this is an anchor for the cost and power overheads (which scale at the same pace) of denser compute and bigger scale-up systems with the same chip. Rubin Ultra Kyber is one step higher in the scale-up and density hierarchy compared to Rubin Oberon, so I'm going to assume that it needs an IT power budget of 1.67 kW per compute die (1.15x more than Rubin), and will cost the same $15.6bn per 1 GW of IT power per year (at multi-year rental prices). This means 600K compute dies per 1 GW of IT power and that the cost is $2.97 per hour of compute die time for Rubin Ultra Kyber . The assumption for Feynman (which I'm keeping from the previous post) is 30% higher IT power budget than Rubin Ultra (the anchor is the IT power budget for Rubin being 1.3x higher compared to GB300), so 2.17 kW per (base) reticle-sized logic die and 460K (base) compute dies per 1 GW of IT power. At the same $15.6bn per year, the cost is $3.87 per hour of Feynman (base) compute die time . The cost for the 2028 model with 240T total params is then $2.5/$46 per 1M input/output tokens , and for the 2031 model with 1.4 quadrillion total params, the cost is $12.3/$67 per 1M input/output tokens . For context, the cost for the 2026 model with 10T total params (that I'm linking to Mythos 5) is $0.79/$12.5 per 1M input/output tokens. These are costs from inference compute (still with a maximum context length of 1M tokens). Future API price estimates then depend on gross margins, and on decisions about maximum context length. Also, more vertical integration that becomes feasible for bigger AI companies can reduce the costs of compute. This post assumes multi-year rental prices, but Hyperscaler TCO for owned compute can be lower by a factor of 1.5x. Improving the models makes inference more valuable, so the amount of training and R&D compute shouldn't be much less than the amount of inference compute. With the revenue that inference compute brings in, a gross margin of 50% pays for about as much training and R&D compute. Gross margins probably can't go below 30% (for the AI company to survive), which probably gets more feasible when growth slows down. The current gross margin of around 85% is not much more useful than a gross margin of 70% (in terms of how much training and R&D compute it enables), but the latter allows doubling the costs per token (at a given price). Total costs (the sum of costs of all tokens served for all requests) are essentially the amount of inference compute, so the current very high gross margins might even be a consequence of inability to access enough compute to increase total costs (which would also happen with lower prices, if demand is not destroyed in other ways), and this problem might get less dire by 2028+ (when bigger models also give a stronger motivation to push the prices somewhat lower). With yet another doubling of the costs, we get a gross margin of 40%, which still seems survivable if the growth (in terms of the amount of compute per AI company) sufficiently slows down by 2031+. The API token usage assumptions from a previous section were 2% output tokens and 80-95% cache hit ratio for input tokens. At output/input token price ratio of 5x and with cache-hit input tokens being 10x cheaper than cache-miss input tokens, the average price per token is 0.24-0.37x the input token price (depending on cache hit ratio). And at the output/input token cost ratio of 16x and with cache-hit input tokens considered free, the average cost per token is 0.37-0.52x the input token cost. The price/cost ratio for an average token is 0.66-0.73x the price/cost ratio for a cache-miss input token (which is relevant for the 2026 model, and for the 2031 model with a maximum context length of 3M). This changes a bit to 0.59-0.67x for the output/input token cost ratio of 18x (relevant for the 2028 model), and changes a lot to 1.2-1.5x for the output/input token cost ratio of 5.5x (relevant for the 2031 model with maximum context length of 1M). At a gross margin of 70% (which I'm going to assume for the 2028 model), the price/cost ratio for an average token is 3.3x, so the price/cost ratio for a cache-miss input token is 5.0-5.6x, and with the cost per 1M cache-miss input tokens of $2.5, the price has to be $12-$14 per 1M input tokens. At the 5x output/input price ratio, the price for output tokens is $62-$70. That is, at a 70% gross margin, the API price for the 2028 model with 240T total params is $14/$70 per 1M input/output tokens . The gross margin for input/output tokens is 82%/34%, so the token usage assumptions are load-bearing. With a maximum context length of 3M (and average request length of 1.5M), the output/input token cost ratio for the 2031 model is 16x, so the price/cost ratio for a cache-miss input token is 4.6-5.1x, and with the cost per 1M cache-miss input tokens of $12.3, the price is $57-$62 per 1M input tokens, and $280-$310 per 1M output tokens. That is, at a 70% gross margin and with maximum context length of 3M, the API price for the 2031 model with 1.4 quadrillion total params is $60/$300 per 1M input/output tokens . The costs with maximum context length of 3M are $12.3/$200, so the gross margin for input/output tokens is 79%/33%. The cheaper option is to keep the contexts at 1M, in which case the output/input token cost ratio is 5.5x and the price/cost ratio for a cache-miss input token is 2.2-2.7x. The price is then $27-$34 per 1M input tokens and $130-$170 per 1M output tokens. That is, at a 70% gross margin and with maximum context length of 1M, the API price for the 2031 model with 1.4 quadrillion total params is $30/$150 per 1M input/output tokens , exactly half the price with the maximum context length of 3M. This time, the gross margin for input/output tokens is 59%/55%, so this is less dependent on the fraction of output tokens in actual use, though going from this to about 70% still relies on a lot of cache-hit input tokens. Finally, the prices that target a gross margin of 40% are simply 2x lower than those that target a gross margin of 70%. In particular, at a 40% gross margin and with maximum context length of 1M, the API price for the 2031 model with 1.4 quadrillion total params is $15/$75 per 1M input/output tokens . The 1M maximum context length option also has the benefit of keeping gross margins for input/output tokens positive, while with the 3M maximum context length, the gross margin for output tokens goes negative, so that the input tokens have to essentially subsidize the output tokens. There is a SemiAnalysis claim that Kyber got delayed, which if correct probably means it's only coming out with first-year Feynman (this also makes it more likely that only second-year Feynman gets 8x Kyber, at least in high volume). But by the end of 2028 there is probably sufficient buildout of alternatives for 240T total param models to be feasible even if there's no Rubin Ultra Kyber. In 2027, TPU 8i finishes its buildout (295 TB per pod, 12-Hi HBM, 3 GB per DRAM die, 1.07 TB/s per stack, so likely HBM3E; 33 ms to fully read). This is a departure from the usual 3D torus topology (which TPU 7x still follows) and a step towards better all-to-all scale-up latency crucial for MoE decode. Though in 2027, maybe only GDM can afford to deploy a flagship model that depends on this system (even Anthropic might want more options for inference). But by 2028 there will be the next TPU, AWS Trainium might have something relevant, and a hypothetical 2-rack scale-up pod based on AMD's Helios upgraded to 12-Hi HBM4E would be almost sufficient (in contrast to how a 2x Oberon will remain insufficient). Nvidia might yet pull through with Rubin Ultra Kyber, or a fallback 4-8x Oberon pod. Thus I'm just going to assume Rubin Ultra Kyber in this post, with possible replacements not changing the costs dramatically, possibly forcing a somewhat lower sparsity on the 2028 model, but keeping its number of active params similar (which is the thing that primarily determines costs). ↩︎ The dimension of activation vectors is also called the model dimension . ↩︎ The sequence of all activation vectors (in different layers) above some token. ↩︎ The paper seems to say it's 1.1 GB for contexts of 32K tokens, see Table 3 , which is 34 KB per token. But the model config says it's 4 KV heads with K=V and dimension 512 in 10 layers out of 60, which is 20K numbers, with sliding window attention adding only 16% for 32K token contexts. Since the discrepancy is modest, I'm keeping the anchor as a range. Incidentally, the model dimension is 5.3K. ↩︎ It's 41B active params. From the config , 11 layers (out of 66) have full attention, 8 KV heads of dimension 128, without K=V. Model dimension is 6.1K. ↩︎ Only shapes of the models that were frontier (largest, most pretraining hungry) can be predicted from pretraining compute. Thus Mythos 5 and Opus 4 qualify, but Opus 5 doesn't (it's very likely overtrained, and probably distilled from Mythos 5). The large jump in Opus 5 capability compared to Opus 4.8 (without a change in prices) supports the hypothesis that the relatively recent Opus 4.8 is still an old pretrain. The tokenizer change in Opus 4.7 is an argument against. A different pretrain doesn't necessarily mean a different model shape (or costs), since hardware targeted by the older pretrain still needs to continue inferencing something (constraining total params), and the range of costs starting from the current frontier model and going down needs to be covered by multiple models (constraining active params). If older models keep shrinking to get cheaper (as it becomes feasible to overtrain them and distill from stronger models), while the shape of the current frontier model is determined by what's compute optimal, this would create a gap in capability that's too large. So arguably smaller models shouldn't shrink (as their pretrains get replaced in a new generation). Instead, new frontier models with more active params should appear on top of the existing models, as continued scaling of pretraining compute motivates their greater size and cost, while retrained models with the older shapes get stronger without getting cheaper. ↩︎ Apart from the trend of how much compute Anthropic had over time, a Dec 2024 Anthropic post claims that the Rainier datacenters being announced there would provide "five times the computing power (in exaflops) used to train our current generation of leading AI models". The reference to the "current generation of leading AI models" is ambiguous, since only Opus 3 was released back then, but Opus 4 was probably already pretrained, so the claim might be referring to it. At that stage in the planned Trainium 2 buildout for Anthropic, the Rainier compute might've referred to 400K chips , which in FP8 would correspond to about 250K H100s (1.3e15 FP8 FLOP/s per Trainium 2 chip compared to 2e15 FP8 FLOP/s per H100 chip). A fifth of that is 50K H100s, 2x less than my assumptions for the big model of 2025 and 4x less than my assumptions for the big 2026 model (which I'm linking to Mythos 5). ↩︎ This is the multi-year rental price until the end of last year, probably when a lot of the GB300 capacity was contracted. The current price is listed as $3.96 per chip-hour in SemiAnalysis's InferenceX (when selecting "Cost per Million Tokens (3 Year Rental)" in the Y-Axis field) (implying $15.6bn per year), and 6-year rental price was stated to be about $4.00 per chip-hour in their Jul 2026 post . ↩︎ The latest public estimate for GB300 from SemiAnalysis is 2,220 W per chip (160 kW per rack), clearly described as before- PUE total IT power (so it includes outside-the-rack networking). There are also estimates for various chips visible in InferenceX (when selecting "Joules per Token" in the Y-Axis field), which are 5% lower for some reason. Bigger datacenters probably have a higher scale-out networking overhead. ↩︎ To make sure everything checks out, let's do the output token cost calculation directly. With 560 KB of KV cache per token and HBM footprint factor of 2.9x, a 500K-long request covers the cost of reading 1.6 TB of HBM per pass (across the whole pipeline). Producing 1M output tokens at 3x MTP speedup then involves reading 540 PB of HBM in 333K passes. At MBU of 70% and HBM bandwidth of 14 TB/s, this takes 15.4 hours of compute die time. ↩︎ Buildout of a post-Feynman Nvidia system is underway over 2031, but I'm stopping at estimating second-year Feynman in this post. ↩︎ The direct calculation starts with 1.4 MB of KV cache per token and HBM footprint factor of 1.6x, so that a 500K-long request needs to cover the cost of reading 2.2 TB of HBM per pass (only 1.4x more than the 1.6 TB for the 2028 model). Producing 1M output tokens at the MTP speedup of 3x in 333K passes then involves reading 740 PB of HBM, which at MBU of 70% and HBM bandwidth of 16.9 TB/s then takes 17.35 hours of (base) compute die time. ↩︎ To see prices for chip-time, select "Cost per Million Tokens (3 Year Rental)" in the Y-Axis field. ↩︎ To see IT power budgets for chips, select "Joules per Token" in the Y-Axis field. ↩︎ Discuss
Score: 40🌐 MovesJul 27, 2026https://www.lesswrong.com/posts/Rk6FbkDFFm8ciqefv/quadrillion-param-costs-kv-cache-context-length-frontier - Run Claude Managed Agents with Chat SDK
You can now run Claude Managed Agents with Chat SDK . Claude Managed Agents handles the agent loop server-side, including the model, tools, session state, and sandboxed web research. Chat SDK gives that agent a chat interface through a single type-safe handler, with adapters that carry it to Slack, WhatsApp, and more. What you get Token-by-token streaming : Replies render as the model writes them, over a single streamed response. Live activity feed : Tool calls and model requests are available as the turn runs, so you can surface a trace in the chat. No database to run: The Managed Agents session stores the conversation, so the sidebar, transcript, and replay read from it, no server-side state of your own. Portability by design : Swapping a few lines in the handler moves your agent to Slack, Teams, Discord, WhatsApp, and 30+ other platforms. To see it work, a new Anthropic quickstart builds a working research analyst you chat with in the browser. It runs on Chat SDK's web adapter, so there's no platform registration, webhook verification, or tunnel, just your Anthropic credentials. Follow the step-by-step guide or learn more in the Chat SDK documentation . Read more
- Sonder brand bought out of bankruptcy by TravelAI
Vancouver-based agentic travel company resurrects Canadian-founded unicorn as affiliate booking site. The post Sonder brand bought out of bankruptcy by TravelAI first appeared on BetaKit .
- How a Pair of 21-Year-Olds Built a $13 Million AI Startup Without Touching Their Funding
Turbo AI’s co-founders launched the company while still undergrads and grew to 10 million users with a team of just 10 people.
Score: 40🌐 MovesJul 27, 2026https://www.inc.com/diana-bocco/student-founders-entrepreneurs-ai-startup-venture-capital-funding/91380311 - A bipartisan House duo’s final pitch for pre-midterm AI action
“There is real urgency,” said Rep. Lori Trahan.
Score: 40🌐 MovesJul 27, 2026https://www.politico.com/news/2026/07/27/ai-house-bill-obernolte-trahan-01012080 - Building the enterprise environment for agentic AI
For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflows, data, and systems. The platform best-suited to run agents is built with proper CPU capacity, resilient data access, policy-aware tool use, observability, memory management, and the…
Score: 40🌐 MovesJul 27, 2026https://www.technologyreview.com/2026/07/27/1140668/building-the-enterprise-environment-for-agentic-ai/ - Wiz And Google Have An Answer To Mythos, And It’s Not A Model
Wiz and Google announce a multi-model system that outperforms Mythos at discovering software vulnerabilities, offering new implications for defenders.
Score: 40🌐 MovesJul 27, 2026https://www.forbes.com/sites/timkeary/2026/07/27/wiz-and-google-have-an-answer-to-mythos-and-its-not-a-model/ - Sam Altman says one of his biggest fears is that a small number of companies will control AI: 'That'd be very, very bad'
Sam Altman says one of his biggest fears is that a small number of companies will control AI: 'That'd be very, very bad' Business Insider
- Big Tech is forcing consumers to pay for its AI boom. Voters are pushing back.
State regulators are making Big Tech pay for its own grid build-out — leaving traditional utility stocks exposed to a political reckoning.
- AI Agent Orchestration For ASIC Autonomy
How to achieve a 10X productivity improvement using AI agents. The post AI Agent Orchestration For ASIC Autonomy appeared first on Semiconductor Engineering .
- ChatGPT Ads adds conversion bidding and stronger measurement
New capabilities bring it closer to the workflow and optimization tools advertisers already use in Google Ads and Meta Ads. The post ChatGPT Ads adds conversion bidding and stronger measurement appeared first on MarTech .
Score: 39🌐 MovesJul 27, 2026https://martech.org/chatgpt-ads-adds-conversion-bidding-and-stronger-measurement/ - HCLTech To Set Up First AI Data Center in Odisha
HCLTech To Set Up First AI Data Center in Odisha india.entrepreneur.com
Score: 39🌐 MovesJul 27, 2026https://india.entrepreneur.com/business-news/hcltech-to-set-up-first-ai-data-center-in-odisha - Orange, Morrison plan French data centre venture to meet AI demand
Orange, Morrison plan French data centre venture to meet AI demand Reuters
- Critical care doctor says patients’ reliance on chatbots reflects deeper problems in health care
Critical care doctor says patients’ reliance on chatbots reflects deeper problems in health care EurekAlert!
- Can AI's power problem become a clean-energy opportunity?
Artificial intelligence has triggered a new gold rush. This time, companies are competing not for land or minerals, but for access to electricity.
Score: 39🌐 MovesJul 27, 2026https://techxplore.com/news/2026-07-ai-power-problem-energy-opportunity.html - Is Mythos good at cyber because it kept hacking Anthropic during training?
From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize the resulting transcripts. [...] The most notable finding was that the model occasionally circumvented network restrictions in its training environment to access the internet and download data that let it shortcut the assigned task —a form of reward hacking. While highly concerning, this behavior was rare, even in settings where it could have been viable and helpful, with attempts appearing in about 0.05% of all training episodes and successful attempts appearing in about 0.01% of episodes. The technique matched the sandbox-escape incident that we separately elicited in our automated behavioral audit when we had an investigator model explicitly ask Claude Mythos Preview to find such a bypass. In every observed instance, the model used this access solely for completing the task. More broadly, we observed the model escalating its access within its execution environment when blocked: reaching a shell from restricted GUI computer-use interfaces, injecting commands through tool-call arguments, or recovering information the task had deliberately hidden. Prompts asking models to avoid this kind of action were not always effective. The recursive-summarization tool flagged this access escalation in about 0.2% of training episodes. By extrapolating from public data (see details below), I estimate that Mythos preview: Escalated its permissions on ~100,000 RL rollouts. [1] Broke sandboxes in ~10,000 RL rollouts (and was likely rewarded for it). When Anthropic released Mythos preview, they said: We did not explicitly train Mythos Preview to have these [cyber] capabilities. Rather, they emerged as a downstream consequence of general improvements in code, reasoning, and autonomy. However, Mythos preview was probably rewarded for hacking Anthropic tens of thousands of times over the course of training . If these hacks involved learning various different cyber strategies, this could've meaningfully increased Mythos's cyber offense abilities. [2] My current guess is that, if Mythos preview did not reward hack at all during training (because the environments were more robust), it would be noticeably less capable at cyber out-of-the-box. However, it would likely be a more generally competent and aligned model, such that a bit of additional cyber training could make it even more powerful than Mythos preview. [3] Thoughts and reflections about this probable fact I feel like they were "sane-washing" these incidents in the Anthropic system card. Anthropic wrote: "the model occasionally circumvented network restrictions in its training environment to access the internet and download data that let it shortcut the assigned task [...] [It succeeded at this] in about 0.01% of episodes." If they instead said: " we estimate that Claude broke its sandbox and accessed the internet tens of thousands of times during training ," I would've thought of this much earlier. Also, everybody fixated on the specific story of Sam Bowman getting an email from Mythos while eating a sandwich in the park because it's so memetic (and fun! haha Claude so cute. sandwiches sandwiches park park). I started thinking about this stuff more after Buck Shlegeris and Ryan Greenblatt said on their podcast about the OpenAI/Huggingface incident that the OpenAI model had probably been repeatedly rewarded for breaking out of sandboxes during training. [4] The more I thought about this, the more true it felt, especially given that we've seen similar behavior before: an early o3 checkpoint would literally go looking for the grader file (see Figure 63 here and Figure 58 here ). Turns out this was explicitly discussed in an Anthropic system card from three months ago. Estimating how many RL rollouts went into Mythos Preview Here are three datapoints that I used: Minimax 2.5, a 230B-A20B Chinese model released in February 2026—the same month that Mythos was deployed internally—was trained on " over a hundred thousand distinct real-world agent scaffolds and environments." I think it's reasonable to expect Mythos to be trained on at least 10x as many environments. If you expect ~10 tasks per RL environment and ~10 rollouts per task, we arrive at an estimate of 10^8 total rollouts (i.e., hundreds of millions). Alternatively, Olmo 3-32B was trained on ~100,000 prompts with a group size of ~10. If we think Mythos preview is trained on around 100 times more prompts than Olmo 3-32B, then this also yields something on the order of 10^8 total rollouts. (Additional datapoint I got after making these estimates, credit to JS at Epoch) Inkling was trained on 30 million RL rollouts . I would naively expect Mythos to be trained on like 3-10x more data than Inkling, which also lands us in the ~10^8 range. Based on the frequencies of misaligned behaviors in the system card, 10^8 total rollouts imply ~10^4 rollouts where it breaks out of sandboxes to download data from the internet and ~10^5 rollouts where it escalated permissions within its execution environment. I suspect that there would also be issues with Opus 4.6 mislabeling transcripts, but I don't expect that to be off by more than a factor of two. ^ I'm only estimating orders of magnitude here, so I'm rounding to the nearest order of magnitude (2×10^5 → 10^5.) ^ The system card says the sandbox-escape technique Mythos used in the behavioral audit matched the one it used during training—though it's unclear whether there was only one. My guess is that Mythos didn't use many distinct sandbox-escape strategies, but did for privilege escalation: broadly applicable escalation techniques tend to get patched once discovered, so repeated success likely required a more varied set. ^ I'm around 70% confident that my story is correct. I'm more confident of the claim: "It's wrong to say that Anthropic 'did not explicitly train Mythos Preview to have [cyber] capabilities' given that their model did all sorts of hacking during training." If I'm wrong, my guess is I'm wrong because cyber capabilities are sufficiently general that doing a bunch of coding training can generalize to strong cyber capabilities regardless of reward hacking. Another potential reason is that the reward hacks the model conducted were pretty straightforward and did not contribute to much learning. ^ Adam Karvonen also speculated in this tweet that Mythos breaking out of sandboxes a ton during training is why it's so good at cyber (responding to John Schulman's tweet on how inoculation prompting made it so that models spent the whole RL run practicing breaking out of sandboxes). I hadn't seen this until after I drafted the post. Maybe others have brought up this idea as well, but I've not heard of it. Discuss
Score: 38🌐 MovesJul 27, 2026https://www.lesswrong.com/posts/QKDoZe6EKhxnFjLWK/is-mythos-good-at-cyber-because-it-kept-hacking-anthropic - AI Sovereignty is Your Alpha: How to Avoid Transferring Your Alpha to a Hosted Model Provider
AI Sovereignty is Your Alpha: How to Avoid Transferring Your Alpha to a Hosted Model Provider Palantir Blog
- Kimi AI and kvcache-ai Open Sources 'AgentENV': A Distributed System that Powers Agentic Reinforcement Learning (RL) Training for Kimi K3
Kimi AI and kvcache-ai Open Sources 'AgentENV': A Distributed System that Powers Agentic Reinforcement Learning (RL) Training for Kimi K3 MarkTechPost
Score: 38🌐 MovesJul 27, 2026https://www.marktechpost.com/2026/07/27/kimi-ai-and-kvcache-ai-open-sources-agentenv/ - Claude Code Prompt Caching: Stop Paying for the Same Context
A practical guide to cache prefixes, token usage, TTLs, invalidation, and the habits that keep long Claude Code sessions efficient. A long coding session can look expensive because the token counter is enormous. But the useful question is not how many tokens appeared in the context. It is how many were processed from scratch—and how many arrived through a warm cache. Claude Code prompt caching can reuse stable context across turns, reducing the cost of repeatedly processing the same prefix. The least interesting way to reduce Claude Code token usage is to shave five words from every request. The more important lever sits underneath the conversation: prompt caching. Claude Code repeatedly sends a large context containing system instructions, tool definitions, project guidance, and conversation history. If the beginning of the request matches an earlier request, Anthropic can reuse that stable prefix instead of processing it again at the full input-token rate. That changes how you should read a large token number. A session that has touched hundreds of millions of cached tokens is not equivalent to one that processed hundreds of millions of new input tokens. A high cached-token count can indicate that Claude Code is reusing an established context rather than rebuilding it on every turn. The better diagnostic is a ratio: cache read ────────────── total input context If cache reads remain high while fresh input stays proportionate to the work being added, the session is behaving efficiently. If cache reads suddenly collapse, look for a prefix change before blaming the model. What Claude Code prompt caching actually reuses Prompt caching is based on exact prefix matching . It does not independently cache each repository file, paragraph, or tool. Claude Code builds a request in a stable order and the cache can reuse the matching beginning. The conceptual sequence is: SYSTEM Claude Code instructions and tool definitions PROJECT CLAUDE.md files, rules, memory, and project context CONVERSATION messages and tool results accumulated in the session NEW TURN the latest request and whatever follows it On the first request, the system processes the context and writes reusable material into the cache. On the next request, the unchanged prefix can be read from that cache while the latest exchange is processed normally. The reusable conversation prefix grows as the session grows. Cache-read input is priced at roughly one-tenth of standard input, although cache creation and model pricing still matter. Anthropic’s Claude Code documentation describes cache reads as billed at approximately 10% of standard input-token cost. That is why a long, coherent session can remain economical even when its nominal context becomes large. Do not flatten the accounting into “every cached token is free.” There are separate measurements: cache_creation_input_tokens records tokens written to a cache. cache_read_input_tokens records reusable tokens read from it. Fresh input, output, and model-specific rates still contribute to usage. The practical objective is therefore not the largest possible context. It is a stable, relevant prefix with a high reuse rate. Cache creation and cache reads tell different stories Cache creation is the write; cache read is the reuse. A healthy workflow watches both instead of treating all input tokens alike. Imagine a project begins with 40,000 tokens of instructions, tool definitions, and repository context. The first turn may need to process and cache that prefix. The second turn can reuse it and add only the incremental conversation. If the prefix remains stable over ten productive turns, the initial write has supported repeated low-cost reads. Now imagine changing the model after turn eight. The new model uses a separate cache, so the existing prefix cannot simply carry over. The next request rebuilds context for that model. This is why token optimization is closer to systems engineering than copy editing. The expensive event is often a transition: warm prefix → configuration change → cold prefix → cache rebuild A high cache-read number is usually a good sign, but it is not a quality metric by itself. You can cheaply reuse an irrelevant conversation for hours. Efficiency needs two tests: Is the prefix being reused? Is the reused prefix still useful for the task? The first saves tokens. The second saves judgment. Stable cache prefixes can improve latency and economics, but only when the retained context still serves the work. How the Claude Code cache grows turn by turn Each turn can reuse the matching prefix and append new conversation, making continuity cheaper than repeatedly reconstructing the same context. A simplified three-turn session looks like this: TURN 1 [system][project][message 1][response 1] processed and cached TURN 2 [cached matching prefix][message 2][response 2] TURN 3 [longer cached matching prefix][message 3][response 3] This explains an apparent paradox: longer sessions can be efficient, while restarting constantly can be wasteful. If each new session reloads the same tool definitions, repository rules, and project context, you keep paying the cache-creation cost. Preserving a relevant session lets the accumulated prefix work for you. But “never start fresh” is equally crude advice. When the task changes completely, stale context can produce wrong assumptions, distract the model, and make verification harder. The right boundary is semantic, not superstitious: Continue when the next task depends on the same repository state and decisions. Use /compact at a natural boundary when you need a shorter working history. Use /clear or a new session when the context is no longer an asset. Claude Code’s current documentation says /compact uses the cached conversation prefix to produce a summary, then begins a shorter conversation cache from that summary. Compaction changes the conversation layer; it is not merely a cosmetic edit. Still, it can be the correct trade when relevance is decaying. The three cache layers—and what can invalidate them Claude Code orders system, project, and conversation context so a stable prefix can be reused across successive requests. Thinking in layers makes cache misses easier to diagnose. System layer This includes Claude Code’s base instructions and tool definitions. Changes near the beginning of the request can invalidate everything after them. According to Anthropic’s prompt-caching guide, connecting or disconnecting an MCP server changes the tool set and invalidates the cache. Upgrading Claude Code can also alter the system prompt or tools. Project layer This includes project context such as CLAUDE.md and related instructions. Claude Code deliberately keeps this layer stable during a running session. An important nuance: editing the root CLAUDE.md file during a session does not immediately rewrite the live prefix. The updated instruction is picked up after /clear, /compact, or a restart. That protects the current cache, but it also means the model may not yet be following the change you just saved. Repository file edits are different. Ordinary code changes do not inherently rewrite the cached project prefix. Tool results can communicate changed state as the conversation proceeds. Conversation layer Messages and tool results accumulate here. /compact replaces the older conversation with a summary; rewind truncates the conversation to an earlier prefix. Skills generally append messages, preserving what came before them. These details matter because “editing any file breaks the cache” is false—and “nothing breaks unless an hour passes” is also false. The one-hour and five-minute TTLs are both real The apparent conflict comes from different access paths. Claude Code subscription sessions request a one-hour cache duration. Anthropic notes that if included plan usage is exhausted and the session moves to extra-usage credits, Claude Code can reduce that duration to five minutes. For API usage—including direct API and supported cloud platforms—the documented default cache lifetime is five minutes. A one-hour cache is available as a separate option, and Claude Code API users can request it through ENABLE_PROMPT_CACHING_1H=1. So the operational rule is not “Claude always caches for one hour.” It is: subscription capacity available → typically requests 1 hour API default → 5 minutes API with 1-hour option → 1 hour subscription using extra credits → may fall back to 5 minutes TTL expiry does not corrupt the conversation. It means the next request may need to create the cache again. If you return after a long break, the first turn can be more expensive; subsequent matching turns can become warm again. That also means there is no universal reason to throw away a useful session after the TTL expires. A fresh session must load context too. Choose based on relevance and correctness, not the emotional sting of one cold request. What reliably breaks a warm Claude Code cache The most useful invalidators to remember are: Switching models. Caches are model-specific. Changing the model requires a separate prefix. Changing MCP connections. The available tool definitions change. Compacting. The conversation layer is replaced by a summary and a new shorter cache begins. Upgrading Claude Code. System instructions or tool definitions can change. Letting the TTL expire. The cached prefix ages out and must be written again. Working from a different directory or machine. Claude Code cache sharing is effectively tied to the environment and working directory. Effort-level changes, by contrast, do not alter the cache key according to the current documentation. Parallel sessions launched from the same directory can sometimes share matching prefixes. Worktrees use different directories, so they do not share that cache. Sequential sessions are more likely to match when the Git state and opening prompt are consistent. This suggests a deeper rule: cache locality follows workflow locality. A practical workflow to save tokens in Claude Code Do not contort the work to worship the cache. Use a few habits that also improve engineering clarity. 1. Put stable guidance before transient instructions Keep durable conventions in CLAUDE.md: commands, architecture boundaries, test expectations, and repository-specific rules. Keep the current task in the conversation. Stable context belongs in the stable prefix. 2. Avoid casual model switching Switch when a different model materially improves the work, not as a reflex between planning and execution. A model change may be worth the cache rebuild; it should simply be intentional. 3. Connect MCP tools before the main run If a task needs a database, browser, issue tracker, or design tool, establish those connections before building a long conversational prefix. Mid-session tool changes move an early part of the request. 4. Compact at decision boundaries Use /compact after a milestone: investigation complete, implementation complete, or one subproblem resolved. Ask the summary to preserve constraints, changed files, failed attempts, commands, and the next verification step. 5. Start fresh when context becomes misleading The costliest token is not the uncached one. It is the cached assumption that causes a bad change. A clean handoff should preserve: goal current repository state decisions and constraints files changed tests run and results open risks next concrete action 6. Measure before inventing rituals Inspect the usage telemetry available to you. Watch for sudden changes in cache reads, then correlate them with model changes, compaction, MCP changes, upgrades, directories, and long idle periods. Token dashboards are most useful when they separate fresh input, output, cache creation, and cache reads instead of presenting one intimidating total. The real optimization target is continuity with boundaries Prompt caching rewards continuity. Good engineering rewards boundaries. Those ideas are not enemies. You want stable instructions, consistent tools, and enough session history to avoid rediscovering the project every ten minutes. You also want explicit task boundaries, compact handoffs, and the courage to discard context that has become noise. The mature Claude Code workflow does both: preserve the prefix while it remains an asset → compact when the history exceeds its usefulness → clear when the task identity changes → verify the new state Saving tokens is a consequence of keeping the system legible. Once you understand the prefix, the token dashboard stops looking like a slot machine and starts looking like observability. Stay curious 🍓 PS: On your next long session, note the time of every model switch, MCP connection, compaction, and restart. Compare those moments with cache-read changes. One afternoon of evidence will teach you more than a month of token folklore. Reference Notes These primary sources support current technical details added during repurposing: Claude Code prompt caching — official explanation of prefix order, invalidation events, TTLs, model switching, directories, compaction, and cache metrics. Claude Code costs — official guidance for monitoring and managing Claude Code token usage and cost. Anthropic API prompt caching — official API behavior, cache lifetime options, pricing concepts, and cache-token fields. Claude Code Prompt Caching: Stop Paying for the Same Context was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities
Last week, OpenAI evaluated two models on an exploit benchmark within an isolated sandbox. Guardrails were reduced for testing, and the models found a vulnerability in their environment, accessed the internet, and reached Hugging Face's production database. No human directed the action, but the breach is a clear example of how much more capable malicious attackers are when equipped with powerful AI models. But defenders have the same tools, and a clear advantage: knowledge of their own codebase. Hacks are initiated from the outside, so the single best defense is finding vulnerabilities from the inside before attackers do. Today we're releasing DeepsecBench , a benchmark that evaluates how well different models find cybersecurity vulnerabilities in application code. For each model the report includes recall, precision, cost, and total time, and combines recall and precision into a single benchmark score. Here is a sample of model performance from the leaderboard: Rank Model Level Score Cost Total time 1 GPT-5.6 Sol xhigh 35.58 $55.98 03:39:00 3 Claude Opus 5 medium 28.36 $31.96 00:47:01 8 Kimi K3 high 17.56 $12.38 01:59:00 10 Grok 4.5 high 15.58 $5.60 01:24:00 We built deepsec to make scanning as easy as possible. Now you can use the benchmark report to build a security scanning program that fits your budget and the complexity of your codebase, choosing the right mix of models to run and how often to run them. How the benchmark works DeepsecBench runs on an open-source codebase at a commit state just before a large number of vulnerabilities were fixed. We selected 50 entry-point files and built a golden set of 231 human-judged findings. Each model's score is a recall-weighted F2 (Score = 100 × 5PR/(4P+R)), weighting recall (R) twice as much as precision (P), because missed vulnerabilities will go unfixed, while false positives don't make your codebase less secure. Findings beyond the golden set are classified by a judge model as real or false, and count for or against the model in the precision measure (recall measures only the 231 known findings). The benchmark is run three times and the data published in the report is the median of the three runs. The construction of the benchmark stays secret. We don't disclose the repository, the commit, the files, or the findings, so there is nothing for models to train against. A model reciting memorized fixes would score near-total recall. Instead, the best run finds 30.7%, and 20 of the 25 runs come in under 20%. The cost of capable analysis is falling When we introduced deepsec , thorough scanning on production codebases required using the most capable models, and they were expensive to run. Frontier models from OpenAI and Anthropic still score highest, but open-weight models and more efficient reasoning options are closing that gap, making comprehensive scanning far more cost-efficient. Today, higher price does not buy proportionally more. Kimi K3 , from Moonshot AI, ranks eighth at a score of 17.56 for $12.38 on the high setting, half the top score for about a fifth of the cost. Grok 4.5 set to high delivers near-Kimi performance for less than half the cost, scoring 15.58 for only $5.60. GPT-5.6 Sol set to medium offers the best score-to-cost balance of the top-performing models, taking fifth place with a score of 25.10 at a cost of $17.95. You can interact with and download these charts on the DeepsecBench page . Anthropic's most capable model, Fable 5, is absent because it declines security work, including defensive tasks. We will add security-enabled versions to the benchmark when they are made available. Rank Model Level Score Recall Precision Issues Cost Total time 1 GPT-5.6 Sol xhigh 35.58 30.7% 96.3% 71/231 $55.98 03:39:00 2 Claude Opus 5 max 32.57 28.1% 88.0% 65/231 $127.93 02:34:00 3 Claude Opus 5 medium 28.36 24.7% 70.5% 57/231 $31.96 00:47:01 4 GPT-5.6 Luna xhigh 26.79 22.9% 81.2% 53/231 $24.59 03:01:00 5 GPT-5.6 Sol medium 25.10 21.2% 94.3% 49/231 $17.95 00:30:52 6 GPT-5.5 xhigh 21.20 17.7% 95.5% 41/231 $43.42 02:22:00 7 GPT-5.6 Terra xhigh 19.21 16.0% 95.3% 37/231 $27.81 01:45:00 8 Kimi K3 high 17.56 14.7% 77.1% 34/231 $12.38 01:59:00 9 Grok 4.5 medium 16.54 13.9% 73.3% 32/231 $11.04 01:37:00 10 Grok 4.5 high 15.58 13.0% 77.8% 30/231 $5.60 01:24:00 11 GPT-5.5 medium 12.63 10.4% 92.3% 24/231 $14.94 00:24:05 12 GPT-5.6 Terra medium 12.12 10.0% 92.0% 23/231 $5.15 00:13:15 13 GPT-5.6 Luna medium 12.08 10.0% 82.8% 23/231 $5.23 00:11:59 14 Gemini 3.6 Flash medium 11.56 9.5% 80.0% 22/231 $10.74 00:26:21 15 Kimi K3 medium 11.48 9.5% 64.1% 22/231 $10.84 01:29:00 16 GLM-5.2 medium 10.46 8.7% 62.5% 20/231 $11.09 01:03:00 17 GLM-5.2 high 9.90 8.2% 53.8% 19/231 $51.84 02:30:00 18 Gemini 3.6 Flash high 8.99 7.4% 77.3% 17/231 $5.33 00:22:18 19 Claude Opus 4.8 max 6.91 5.6% 81.3% 13/231 $7.56 00:19:29 20 Claude Opus 4.8 medium 6.88 5.6% 61.9% 13/231 $10.81 00:18:16 21 Inkling high 6.32 5.2% 48.0% 12/231 $1.93 00:11:34 22 Inkling medium 5.26 4.3% 37.0% 10/231 $2.18 00:10:31 23 Claude Haiku 4.5 default 4.73 3.9% 33.3% 9/231 $5.91 00:29:40 24 Gemini 3.5 Flash Lite medium 2.14 1.7% 44.4% 4/231 $0.58 00:01:46 25 Gemini 3.5 Flash Lite high 1.07 0.9% 28.6% 2/231 $0.18 00:01:07 The costs reported in the benchmark cover 50 files. A production codebase can run on the order of 100 times that, pricing a full pass at roughly $1,200 for a Kimi K3 sweep, or over $5,000 for the top-scoring frontier model from OpenAI. Building a scanning program across multiple models Because security scanning is recurring work, the question is which model to use for which task, and when. For example, frontier models can be used for periodic deep audits, while cheaper models like Kimi K3 or Grok 4.5 run scans at a higher cadence. Turning GPT-5.6 Sol's reasoning down from xhigh to medium moves it from 35.58 in 3 hours, 39 minutes to 25.10 in just over 30 minutes, fast enough for reviewing new features before they are pushed to production. A startup might scan every merge with Grok 4.5 and save more expensive audits for milestone releases. A large enterprise might run frontier audits on its critical services, the same model at lower reasoning on key pull requests, and a continuous Kimi-class or Grok sweep across the rest of the codebase. Running the benchmark on one endpoint Security scans are spiky workloads that consume a large number of tokens in a short time window, and that burst capacity is hard to buy from any single provider directly. Every DeepsecBench run goes through AI Gateway . The gateway gives us a single endpoint to reach each model. During runs, routing, retries, and failover happen automatically, without any per-provider keys or rate limits to manage. The same routing we use for the benchmark can run your security scanning program. Pass provider/model to deepsec , and one AI_GATEWAY_API_KEY in your environment covers every model on the board (or on a linked Vercel project, access AI Gateway via OIDC). Find your vulnerabilities first Attackers can use AI models, but they act from the outside. Defenders can see their entire system: the source, the architecture, the history. That visibility makes the same models more powerful in your hands than in theirs, because a model that can read the source finds what an attacker can only probe for. The advantage is real, but it only counts if you use it first. We will continue to update DeepsecBench with new models as they are released and measured. Read more
Score: 38🌐 MovesJul 27, 2026https://vercel.com/blog/deepsecbench-evaluating-model-performance-in-finding-cybersecurity-vulnerabilities - Off-grid data centers aren't the panacea that some AI firms hoped
Local backlash and reliability concerns are challenging the vision of an AI boom built quickly using large off-grid data centers. Why it matters: If giant projects stumble, trillions of dollars in AI investment and the industry's plans to rapidly expand data centers could be at risk. The big picture: In an effort to speed AI development, some companies are building data centers powered mainly by onsite generation rather than waiting years to connect to the electric grid. There are 59 data centers with a combined capacity of about 90 gigawatts that plan to build their own power "behind-the-meter" using sources like gas turbines, generators and fuel cells, according to a report from research firm Cleanview. A smaller subset is trying to run large campuses mostly on behind-the-meter power as a way to move more quickly than the local utility and the grid can. Infrastructure risk analysis firm Occam Edge tracks 12 projects where on-site power is the primary serving supply, representing about 10.6 GW of announced capacity. State of play: Recent hiccups are highlighting the challenges that come with bypassing the grid. Earlier this month, New Mexico's top land official rejected a gas pipeline meant to supply on-site fuel cells for Oracle's 2.5 GW Project Jupiter data center campus, part of Oracle and OpenAI's Stargate initiative. The regulatory setback could contribute to a yearslong delay . A week later, a much smaller off-grid data center in Virginia , saw its on-site gas turbines knocked offline for 24 hours , forcing it to run on dirty backup diesel generators during already bad air quality from the Canadian wildfires. Local residents complained of burning lungs and 60-decibel noise, and a local government official called for new laws to regulate generators. Catch up quick: Chief developers of this off-grid approach include OpenAI, and its partners Oracle and Crusoe, which have worked on OpenAI's Stargate campuses. A Stargate project in Abilene, Texas, developed by Crusoe, reportedly went offline for days at a time due to issues with power and cooling equipment. Another off-grid player is Elon Musk, who rattled the energy world when he built xAI's Colossus 1 using mobile gas turbines in just a few months. While it was initially off-grid, it's now grid-connected and selling compute to Anthropic. Despite a lawsuit from the NAACP over Colossus 1, Musk is now using gas turbines to power Colossus 2 (with compute sold to Google), and earlier this month bought a mobile gas turbine company. What they're saying: Critics argue off-grid data centers could prove slower and more costly to deploy, less reliable to run, and more expensive than expanding the electric grid. "This is a flimsy way to deploy AI," says energy investor Jigar Shah, who predicts much of the bullish off-grid deployment figures won't materialize. Power engineers are worried that off-grid data centers won't meet reliability targets, says Occam Edge founder and CEO Christian Okoye, who compares "Dark Gigawatts" to the "Dark Fiber" of the dotcom bust. "It's a time of desperation," says Josh Wong, founder and CEO of ThinkLabs AI, which helps utilities find more capacity on their current grids. "The [off-grid] business model is bolt-on." Follow the money: Investors are lacking insight into the reliability of off-grid data centers, says Okoye. Investors faced with the uncertainties of off-grid data centers could decide not to underwrite them. S&P Global Ratings recently downgraded Oracle's long-term issuer credit rating to BBB- from BBB, just one step above junk status, due to its massive data center spending, which includes investment in on-site power infrastructure. Trillions of dollars are riding on whether large off-grid data centers can work as advertised, justify their added cost and win local support. What's next: More high-profile problems building and running off-grid data centers could lead to a reckoning that the AI boom will be a lot more grid-powered than many AI companies hoped.
- AI can fuel biological weapons. We must harness its power for defense | Annie Jacobsen
The offensive potential is no longer theoretical. We need to develop systems to strengthen public health as quickly as AI is accelerating biological design As artificial intelligence rapidly transforms the biological sciences, it is pushing the future of biology in two opposing directions. AI can help bad actors generate recipes for biological weapons with just a few keystrokes and computational prompts. At the same time, AI can track disease outbreaks and deliver critical public health information to millions of people in real time – potentially stopping a deadly outbreak in its tracks. What we are dealing with in this moment is profound: a race between offense and defense. Between the proliferation of dark biology, where deadly pathogens are engineered in secret, and the promise of a new era of collective global health. Every infectious disease outbreak begins with a handful of cases. Public health wins by learning about an outbreak before it spreads wide, before hospitals fill and a crisis starts to spiral out of control. As transmission rates explode, options shrink. One of AI’s greatest values is its ability to compress the timeline between outbreak and detection. Speed can mean the difference between containing an outbreak and confronting an epidemic, or another global pandemic. Continue reading...
Score: 38🌐 MovesJul 27, 2026https://www.theguardian.com/commentisfree/2026/jul/27/ai-biological-weapons-defense - AI chatbots citing Russian propaganda sourced from EU-sanctioned outlet
AI chatbots are citing Russia propaganda from EU-sanctioned outlets, researchers say, raising questions about how hybrid warfare is expanding to influence large language models (LLMs).
- Artist sues AI meme generator for selling deeply personal comic as ad template
Meme generator may have screwed up by using templates in outputs, expert says.
- Cadence raises annual forecasts as demand booms for AI chip design
Cadence raises annual forecasts as demand booms for AI chip design Reuters
- JPMorgan’s AI leadership is getting a reset. These are the key players.
JPMorgan’s AI leadership is getting a reset. These are the key players. Business Insider
Score: 38🌐 MovesJul 27, 2026https://www.businessinsider.com/jpmorgan-ai-leadership-teresa-heitsenrether-retire-jamie-dimon-2026-7 - Building a product company from Tamil Nadu; LiteFold is accelerating drug discovery
Building a product company from Tamil Nadu; LiteFold is accelerating drug discovery YourStory.com
Score: 38🌐 MovesJul 27, 2026https://yourstory.com/2026/07/building-a-product-company-from-tamil-nadu - HNS 2026 | Huawei Unveils Upgraded Xinghe AI Campus Solution for Southern Africa
HNS 2026 | Huawei Unveils Upgraded Xinghe AI Campus Solution for Southern Africa The Straits Times
- The Next Evolution of Robotics: From Standalone Machines to Coordinated Autonomous Workforces
The Next Evolution of Robotics: From Standalone Machines to Coordinated Autonomous Workforces Toronto Star
- Patients warned off using AI chatbots for self-diagnosis as flaws revealed
Patients warned off using AI chatbots for self-diagnosis as flaws revealed thenationalnews.com