AI News Archive: July 21, 2026 — Part 15
Sourced from 500+ daily AI sources, scored by relevance.
- Advancing next-gen AI with materials science innovation
The conversation about AI often centers on algorithms, computing power, or huge investments in new semiconductor fabrication plants and hyperscale data centers. But beneath each of these advances is another layer of innovation that makes them possible: advanced materials. Every new generation of AI technology demands more processing power, more memory, greater energy efficiency, and…
- Why AI Needs a “Genie Coefficient”
Proposing a new metric for whether AI does what you actually want
- RIL Q1 FY27: JioStar launches AI-generated micro-drama, brings ChatGPT to JioHotstar search
JioStar has launched its first fully AI-generated micro-drama, Game On: 4,000 Crore Empire, on the Tadka platform and integrated OpenAI’s ChatGPT into JioHotstar’s voice search, signalling a major AI push across content creation and discovery. The post RIL Q1 FY27: JioStar launches AI-generated micro-drama, brings ChatGPT to JioHotstar search appeared first on MEDIANAMA .
- Nvidia Touts Progress Getting New Rubin Design to Customers
Nvidia Corp., the semiconductor company at the heart of the artificial intelligence boom, said its latest chip designs are making their way to customers and will help solidify the chipmaker’s leadership in the industry.
- Anthropic’s Robot Ambition; Nvidia Ramps Up Vera Rubin
Anthropic’s Robot Ambition; Nvidia Ramps Up Vera Rubin The Information
- Nvidia is releasing full specs for its AI server CPU, taking aim at AMD and Intel
The chip maker published a white paper and SPEC CPU 2026 results showing Vera ahead of AMD's Epyc 9755 in integer performance
- NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI
Agentic AI shifts more of the critical execution path onto the CPU. Agents operate in sandboxes to execute code, invoke tools, retrieve context, interact with...
- Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token...
- NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
NVIDIA Vera Rubin is here, and it’s going gigascale. Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Spanning 350+ factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand. The […]
- Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories
AI has entered the gigascale era. The world’s most advanced AI factories are bringing together hundreds of thousands of GPUs and CPUs to train frontier models, power agentic AI and generate intelligence at unprecedented scale. At this level, networking becomes a critical computing power multiplier in driving token generation. Marking a networking milestone, NVIDIA Spectrum-6 […]
- Behind the scenes at Nvidia's Engineering SuperLab — Vera Rubin NVL72 running OpenAI workloads, 800VDC demonstrated, and more
Nvidia gave Tom’s Hardware an exclusive look inside its previously undisclosed Engineering SuperLab near Nvidia HQ, where we saw Vera Rubin NVL72 in action.
- Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack
Nvidia has detailed new features of its Rubin architecture.
- Nvidia has shipped 'hundreds of thousands of Grace standalone servers’ — GPU firm pivots messaging as CPUs take center stage in agentic data centers
As Nvidia continues to roll out Vera, its first custom CPU for agentic AI, it revealed that its last-gen Grace design has seen mass deployments, even as a standalone CPU for non-agentic workloads.
- Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains
Artificial intelligence chip king Nvidia Corp. today revealed a fresh trove of performance benchmarks and architectural milestones for its next-generation Vera Rubin platform as it edges closer to global availability. The new numbers are impressive, but on a higher level they also underscore the potency of Nvidia’s approach that combines custom silicon with its broader […] The post Nvidia doubles down on AI factories as it showcases massive Vera Rubin performance gains appeared first on SiliconANGLE .
- Author Invited to Give Speech at OpenAI Headquarters, Uses Opportunity to Trash AI to Their Faces
"I had a ball." The post Author Invited to Give Speech at OpenAI Headquarters, Uses Opportunity to Trash AI to Their Faces appeared first on Futurism .
- Apple prepares a new way to buy iPhones. Plus, making sense of Google's new AI models
Every weekday, the Investing Club releases the Homestretch; an actionable afternoon update just in time for the last hour of trading.
- Last Week in AI #251 - Mythos Back, Sonnet 5, Etched, LongCat
Trump lifts restrictions on Anthropic, Anthropic launches Claude Sonnet 5, Google's NotebookLM updates, chips stories from Etched and Baidu, and more!
- Last Week in AI #250 - Mythos Mess, GPT 5.6-Sol, GLM 5.2
Anthropic's AI treaty discussions, US government's influence on AI model releases, OpenAI's processor development, memory market impacts, and more!
- OpenAI and Anthropic are breaking their own lobbying records as IPOs loom
Together the two AI companies spent $3.17 million on federal lobbying in the second quarter, up 23% from the first quarter
- Opinion | Gov. Kathy Hochul: How New York Will Get AI Data Centers Right
I want AI companies to succeed here.
- Google is building an AI fence around the internet it once championed
Google is building an AI fence around the internet it once championed
- The Debate About Chinese Open-Source AI; Ellison’s Bad Day
The Debate About Chinese Open-Source AI; Ellison’s Bad Day The Information
- Chinese open-weight models are cheap. Washington is deciding what that costs.
Enterprises evaluating Chinese open-weight models this month face a question that has nothing to do with benchmarks: whether using one will still be straightforward in a year. Moonshot AI’s Kimi K3 arrived on July 16 as the largest open-weight model yet released, and within days it had reopened a policy argument in Washington that had […] The post Chinese open-weight models are cheap. Washington is deciding what that costs. appeared first on AI News .
- Lets Be Realistic About Kimi Open Source
This article was written after I saw an interview from the DataBricks CEO, which suprised me a lot. Kimi K3 has earned the excitement… Continue reading on Towards AI »
- How Texas leaders are responding to growing data center backlash
How Texas leaders are responding to growing data center backlash Houston Chronicle
- Are the US and China entering an AI Cold War?
Are the US and China entering an AI Cold War? thenationalnews.com
- Yet another study says AI is bad for elections, and the rabbit hole gets worse
ChatGPT and Gemini failed to reliably match voters with Hungarian parties, adding another warning as campaigns and manipulators learn how easily AI answers can be shaped.
- Fabricated AI animal videos are distorting how we see wildlife, researchers warn
AI-generated animal footage is racking up millions of views online, but researchers warn it is also distorting public understanding of wildlife behaviour and could be undermining support for conservation.
- The Proliferation of Deepfakes Is Profoundly Changing the Way Teenagers Use the Internet
"We grew as AI grew. We grew up together. Adults don't have the same life experience that we do. They don't understand how advanced and how intense AI is." The post The Proliferation of Deepfakes Is Profoundly Changing the Way Teenagers Use the Internet appeared first on Futurism .
- The AI Copyright Lawsuits Have Finally Produced an Actual Payout
Anthropic’s massive $1.5 billion settlement just received final approval.
- Could 'banned topics' be the secret to stopping AI hackers?
The trick to stopping an AI cyberattack may be as simple as bringing up a subject the AI system is not allowed to discuss.
- Using LLMs to Find and Prioritize Vulnerabilities Is No Easy Task
The latest large language models have high false-positive rates and fail to take into account the context of scans, leading to more work for AppSec professionals.
- Choose Wisely: AI-Generated Coding Risk Varies, a Lot
AI-generated code introduces 15 vulnerabilities on average per codebase, but the actual risk depends on framework pairing more than the model used.
- Indian clinicians trust AI more than global peers, but use it less at work
A report by Elsevier found that Indian clinicians face relatively lower workplace pressures than their global peers, as nearly 79 per cent said they had sufficient time to provide good patient care
- WAIC 2026 Robotics Exposed: Four Fundamental Changes After 30K Steps Through the Exhibition Hall
After walking 30,000 steps across WAIC 2026 robot exhibition hall, four shifts emerge: from spectators to buyers, motion control democratization, universal to vertical scenarios, and embodied AI moving from demo to delivery.
- When Your Agents Go Dark: Observability in Multi-Agent Systems with OpenTelemetry
Written by Matteo Rossi . In recent months, AI applications have radically evolved. Earlier, prototypes looked like a single loop: prompt, model call, optional tool call, response. In production, an agent often becomes a small system: model calls, tool calls, routing logic, memory, retries, and error handling. We can have an agent that acts as a router: delegating tasks to other specialized agents; a retrieval agent that fetches information from databases and sends it to an agent responsible for reasoning; or an agent that critically reviews the work of other agents. Each of these agents will have its own model calls, its own reasoning and execution logic, and at each of these points, a potential point of failure and an error-handling mechanism. The parallel with traditional systems is more applicable today than ever. Once two agents coordinate through a handoff, a queue, a tool call, or shared state, the system has distributed-system failure modes. That means the debugging problem changes. The hard part is no longer fixing the failing step; it is finding which step failed and why. Years of production incidents have given us an operational model built on logs, metrics, traces, and alerts tied to expected service behavior. For new AI applications that use agents, we often have the same tools, but not the trace structure (or attributes) needed to answer the question: what happened and why? Logs can show local events, but they do not show the request path unless those events are correlated into a trace. The standard for producing those correlated traces is OpenTelemetry (OTel): an open-source, vendor-neutral observability framework hosted by the Cloud Native Computing Foundation that defines a common way to generate and export telemetry. It gives us a protocol for shipping data (OTLP) and a shared vocabulary of attribute names and semantic conventions, so that a span means the same thing across tools. The rest of this article applies it to multi-agent systems. This article walks that path in order: the failure modes that logs miss, how to instrument an agent handoff in code, what the resulting trace looks like in a real run, and when the simpler approach is still the right one. How We Got Here Let’s take it one step at a time to understand what we’ve done so far and the main issues with this approach. This is essential for us to understand how we can move toward full observability . The first agent that most teams develop consists of a single loop: a prompt, a call to a model, a call to a tool (if needed), and a response. Debugging in this scenario is straightforward: log the prompt, log the calls to the tools and the model, and finally, the response. If something goes wrong along the way, we can review all the logs linearly and figure out where the error is. Everything lives within the same process, like a traditional monolithic application. For an agent this simple, structured logging is enough, and the tracing infrastructure described later would add cost with little return. The system grows and becomes more complex. Instead of a single, all-purpose agent, we now have a router that delegates tasks to specialized agents. In addition to the various specialized agents, we need to add a memory management service to persist the context and an orchestrator to manage everything consistently. Figure 1. A representative multi-agent orchestration. One request spans five services, and without a shared trace ID, their logs cannot be correlated into a single request path. This is a representative example of a common orchestrator pattern, where a router delegates to specialized agents, rather than a specific product. In a scenario like this, each box shown in the diagram logs to its own standard output, and execution is no longer linear. Sometimes control passes to an agent that returns its response. At other times, two or more agents may work in parallel. The earlier mental model of reading logs backward no longer holds because the request path crosses agent boundaries and may branch. Four Problems That Compound The Correlation Problem When five or more agents each generate their own logs, we end up with a timeline of uncorrelated events. Reconstructing the path of a single user request means stitching events across services, comparing timestamps, and guessing which call belongs to which step. Parallel calls make timestamps unreliable, too. This holds whether a person or an automated, agentic log-analysis tool performs the reading. In both cases, the events lack a trace ID, so any reader must infer the request path. The cost of this shortcoming can be measured using one of the DORA (DevOps Research and Assessment) metrics: Mean Time To Resolve (MTTR). A small diagnosis task can become a multi-hour reconstruction exercise when most of the effort is spent piecing together and connecting the dots of what happened rather than resolving the problem. 2. The Causality Problem Even when we find the logs relevant to the incident, we only have part of the story. We know what happened, but we don’t always know why. We can see that an agent returns a particular response. Still, we cannot determine if what we receive is generated by a truncated context window, a retrieved document that was actually deprecated, or a threshold set too generously, allowing the model to hallucinate. Because execution paths can vary, traces need to capture the decisions made along the way: tool choice, arguments, retrieved sources, model parameters, and token use. The same input can produce two different execution paths, depending on what the agent decides to do in a fully autonomous, non-deterministic manner. 3. The Cost Problem The cost principle for an agent is very simple: every agent interaction is a model call, and every model call consumes tokens. In a single-agent system, calculating the cost is straightforward because there is a single stream of requests associated with a single cost account. In a multi-agent system, however, a single user request can trigger an indefinite number of agent actions, with subsequent model calls, each of which may affect the responses provided by the LLM and potentially trigger further calls. A common pattern amplifies this: in an agentic evaluation loop, an agent reviews its own answer (or the evidence supporting it) and retries until it is satisfied. The result is a multitude of model calls that consume a high number of tokens. Attributing that cost back to the agent that generated it becomes difficult. 4. The latency problem Latency is another concern. In multi-agent systems, there are numerous potential sources of delay: serialization/deserialization of responses or calls, a slow-responding model, or a poorly configured timeout. When responses slow down, span timing shows whether time is spent on retrieval, model calls, tool execution, serialization, or waiting on another agent. Tracing Agent Handoffs with OpenTelemetry We have seen a version of this before. It is the microservices observability problem in new clothing. That is good news because distributed tracing provides a useful base, and OpenTelemetry is the standard way to model and export those traces. Agent orchestration adds GenAI-specific attributes on top: A trace represents a user request from the moment the prompt is typed until the agent’s final response. A span represents a single unit of work, such as invoking an agent, calling a tool, retrieving information, or calling a model. Each span is embedded in a parent-child tree, creating a hierarchical relationship that starts with the orchestrator’s span and extends down to the individual operation performed by specialized agents. Context propagation carries the trace ID across process and service boundaries. This allows spans to be joined into a single request trace. OpenTelemetry now includes semantic conventions for Generative AI . It records model calls, token usage, prompts, completions, tools, and related events. This standardized vocabulary includes various attributes such as the model name, the number of tokens used, and the type of operation performed. That keeps the instrumentation portable across open-source and proprietary backends. Figure 2. Proposed instrumentation: one correlated trace per request Each agent emits a span under the orchestrator span. The application exports them over OTLP to an OpenTelemetry collector that forwards traces to a trace backend and metrics to a metrics backend. The trace backend reassembles the spans into a single trace per request, so the full causal path is visible in one place. Instrumenting an Agent Handoff Building on those tracing concepts, we can now instrument an actual agent handoff in code. The principle behind this strategy is straightforward: wrap every meaningful unit of work within a span, and track the relevant information about that specific unit within the span’s attributes. Let’s illustrate this situation with an example, using an agent written in Java and instrumented with the OpenTelemetry APIs. In this example, the agent is delegating the task to a specialized agent. Span orchestratorSpan = tracer.spanBuilder("agent.orchestrator") .setSpanKind(SpanKind.INTERNAL) .startSpan(); try (Scope scope = orchestratorSpan.makeCurrent()) { // Child span for the model call, following GenAI semantic conventions Span policySpan = tracer.spanBuilder("chat policy-agent") .setSpanKind(SpanKind.CLIENT) .setAttribute("gen_ai.operation.name", "chat") .setAttribute("gen_ai.system", "openai") .setAttribute("gen_ai.request.model", "gpt-4o") .setAttribute("gen_ai.request.temperature", 0.2) .startSpan(); try (Scope policyScope = policySpan.makeCurrent()) { PolicyResult result = policyAgent.evaluate(context); // Record the decision and the cost policySpan.setAttribute("gen_ai.usage.input_tokens", result.promptTokens()); policySpan.setAttribute("gen_ai.usage.output_tokens", result.completionTokens()); policySpan.setAttribute("agent.policy.version", result.policyVersion()); policySpan.setAttribute("agent.retrieval.doc_ids", result.sourceDocIds()); } catch (Exception e) { policySpan.recordException(e); policySpan.setStatus(StatusCode.ERROR); throw e; } finally { policySpan.end(); } } finally { orchestratorSpan.end(); } From this example, we can see that the child span is created while the orchestrator’s current span is active, which allows OpenTelemetry to link the two spans within the same trace. A note on versions: the OpenTelemetry semantic conventions for Generative AI are still in Development status and may change between releases. The attribute names used here follow version 1.42, pinned in the companion repository. Confirm the current version before reusing these names. Some attributes are still being renamed or relocated. What to Capture on Each Span It is not necessary to record every attribute included in the convention, though it is worth maintaining a minimal subset that serves as a baseline. In this way, we can transform diagnostic traces into something more meaningful and useful for our purposes. Many of these attributes do not need to be set manually when using Spring AI . Spring AI is the Spring team’s framework for building AI applications in Java that wraps model providers with a common set of abstractions. For observability, it builds on Micrometer, the metrics and tracing facade used across the Spring ecosystem. Micrometer bridges to OpenTelemetry, so the data it produces can be exported over OTLP. On Spring AI 1.0, add two dependencies to the pom.xml ( spring-boot-starter-actuator) plus a Micrometer-OTel bridge, such as micrometer-tracing-bridge-otel . The framework records Micrometer observations for supported model calls ( gen_ai.client.operation ) and tool invocations ( spring.ai.tool ), along with token-usage metrics (gen_ai.client.token.usage ), without manually wrapping each call in tracing code. Those observations become spans and metrics through Micrometer’s bridge to OpenTelemetry. The manual spans created above, on the other hand, serve to show the framework what it cannot see: the orchestration layer. Spring AI does not know the decisions the orchestrator makes between model calls. It doesn’t know which routing logic was chosen or why a specific agent was called rather than another. These decisions lie within the application’s business logic and not in the framework, so they require explicit instrumentation. Observability In Action: A Real Trace Jaeger is an open-source distributed tracing backend, hosted by the Cloud Native Computing Foundation. It collects spans and displays them as a trace tree. We use it here to see what the instrumentation produces. We export the agent’s spans to Jaeger over OpenTelemetry and run a sample orchestration. Any OpenTelemetry-compatible backend would work the same way. Jaeger is convenient because it is open-source and quick to run locally. Figure 3. A single orchestration request as one correlated trace in Jaeger: the orchestrator span fans out to the retrieval, policy, and summarizer agents, and the selected model-call span carries the GenAI attributes, such as model name, provider, and token usage Let’s see how a single orchestration trigger call gives rise to multiple spans, some of which contain calls to models (in this case, Claude Haiku). The left column shows the trace tree. At the beginning is the incoming HTTP call, followed by the span related to the agent.orchestrator . It’s split into three child spans: chat-retrieval-agent , chat-policy-agent , and chat-summarizer-agent . Each of these is further split by Spring AI’s auto-instrumentation into a spring_ai_client call and an outbound call to the model. The timing bar shows the sequence of operations: retrieval, then policy application, then summarization. Within each trace, we can view the captured attributes according to OpenTelemetry’s conventional semantics. Here, gen_ai.system identifies Anthropic as the provider, while gen_ai.request.model and gen_ai.response.model give the requested and returned model names, and gen_ai.usage.input_tokens (140) along with gen_ai.usage.output_tokens (59) report the cost per span. Finally, gen_ai.response.finish_reason indicates that the call completed successfully without errors. This trace shows which agent spent the tokens, which documents were retrieved to inform the policy decision, and where latency accumulated. When The Classic Approach Still Makes Sense This setup is not free, and it introduces dependencies: a collector to install, manage, and update; a backend to store traces; and a small amount of latency and compute overhead within each instrumented agent. As noted earlier, a single agent with a single prompt call and a couple of tools is well served by structured logging, and the setup described here is not worth its cost there. The same reasoning applies to internal prototypes with few users operating in purely experimental contexts. The issue is always complexity and the number of moving parts. The goal is to have observability strategies suited to the context, not ones that are complex because best practices dictate so. The Alternative: Build on a Standard or Buy a Platform OpenTelemetry is not an observability product, but rather a protocol and a vocabulary. The next decision is whether to build the tracing layer ourselves or buy an observability platform. The market is currently full of observability platforms that now support observability for interactions with LLMs. To make informed choices, it is necessary to understand the fundamental difference between value and vendor lock-in for each tool. All of the platforms below interoperate with OpenTelemetry. What differs is how each engages with it; that difference shapes the lock-in. LangSmith , from the LangChain team, is the most ecosystem-native option. Within LangChain/LangGraph, it captures traces, prompts, and token usage, and adds built-in evaluation tools. Its OpenTelemetry support is a bridge layered on a proprietary SDK, since the native instrumentation is LangChain-specific, and OTLP import and export were added later. The trade-off is the dependency on a proprietary hosted platform, whose zero-setup convenience pays off mainly if we stay inside the LangChain ecosystem. Langfuse occupies a middle ground. It is an open-source, self-hostable product focused on LLM tracing, prompt management, and prompt evaluation. It engages with OpenTelemetry natively on the ingestion side. It accepts OTLP spans directly, so that spans instrumented with OpenTelemetry flow in without a separate SDK. This is a pragmatic choice when we want a graphical interface for LLM interactions, while staying on OpenTelemetry standards. Arize Phoenix is designed for evaluation and quality monitoring, including drift detection, hallucination detection, and embedding analysis. It is built directly on OpenTelemetry conventions via OpenInference, a set of OTel semantic conventions and auto-instrumentation packages for LLM applications. Therefore, its traces are standard OTLP. It fits when the goal is monitoring interaction quality rather than operational observability. General-purpose OTel backends ( Grafana Tempo , Jaeger , Honeycomb , Datadog , New Relic ) consume OTLP like any other trace, with no LLM-specific UI. Agent traces sit alongside everything else traversing the infrastructure, such as databases, queues, and HTTP services. The table below summarizes all of these aspects. The underlying principle remains the same in every case. Choose the preferred platform, but ensure instrumentation follows open standards. If the spans that make up the trace follow conventional GenAI semantics and export metrics via OTLP, we can switch backends, change LLMs, or add specific tools to ensure observability without rewriting the core instrumentation model. The shared vocabulary is what makes this work. Once spans follow the GenAI conventions, any OTLP backend reads them identically. Conclusion A multi-agent system has many of the same observability problems as a distributed system, with the added variability of model outputs, tool behavior, and dynamic routing. These systems operate according to fan-out strategies that are unknown in advance and make decisions based on the outputs they receive. All of this makes it practically impossible to rely solely on a log for observability, no matter how well-formatted or well-written it is. Therefore, the goal is to address this problem by collecting traces and metrics and storing them in tools that use a universal vocabulary. Whether we choose an open-source backend or a complete SaaS platform, the principle to apply remains the same. Capture all relevant traces and maintain an open, interoperable vocabulary and standard. The accompanying Java sample shows the trace setup, custom orchestration spans, and exported Jaeger trace. When Your Agents Go Dark: Observability in Multi-Agent Systems with OpenTelemetry was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- The agent problem nobody budgeted for
Why organizations need governance for AI agents
- Claude cracks 87-year-old math problem
Claude AI solves a long-standing mathematical challenge, showcasing advanced reasoning capabilities.
- Could Sparse Attention become the future of generative AI?
Could Sparse Attention become the future of generative AI? verdict.co.uk
- AI that acts, plans, and executes: Understanding the power of Agentic OS
By Sanchit Sood, Chief AI Officer at Kapture CX For the past two years, most conversations about artificial intelligence in the enterprise have revolved around a single idea: assistance. AI […] The post AI that acts, plans, and executes: Understanding the power of Agentic OS appeared first on Express Computer .
- Tribuna.com Goes Inside FIFA's AI-Ready World Cup With Lenovo Ukraine's General Manager
Tribuna.com Goes Inside FIFA's AI-Ready World Cup With Lenovo Ukraine's General Manager azcentral.com and The Arizona Republic
- I tested iOS 26 Siri against iOS 27 Siri AI in my car - and it wasn't even close
With the iOS 27 public beta on my iPhone, I challenged Siri AI to assist me in the car. Here's how it performed over regular Siri.
- AWS standardizes more AI billing data to simplify cost analysis
AWS has updated AWS Data Exports, its service for generating and managing cost and usage Reports (CURs), to include standardized Amazon Bedrock product metadata, making it easier for enterprise engineering teams to analyze AI usage and spending as they scale AI deployments spanning multiple foundation models. The update extends billing exports with normalized fields for model provider, model name, inference type, inference mode, billing unit and Bedrock product family, and will enable enterprises to identify which models generated costs and compare spending across providers without relying on custom parsing or normalization of billing records, AWS wrote in a blog post . That reduced reliance on custom parsing will reduce the engineering effort required to analyze billing data, analysts said. “Before the update, a data engineer would typically need to maintain a model ID registry, write regex against usage type strings, or join AWS CloudTrail with CUR to figure out which provider generated which cost,” said Bhupendra Chopra , chief revenue officer at IT consulting firm Kanerika. The new standardized fields “can be the difference between a billing pipeline that needs constant babysitting and one that doesn’t,” Chopra added. That’s because custom parsing logic is more prone to break down or require maintenance when AWS adds new models or updates pricing in Bedrock, said Pareekh Jain , principal analyst at Pareekh Consulting. Richer billing data to boost enterprise AI cost governance Beyond reducing engineering overhead, the update could also help enterprises improve AI cost governance. Before this update FinOps teams struggled to identify which model Bedrock related to because usage type fields were inconsistent, and there was no unified product family name that captured all Bedrock costs in one place, Chopra said. “Now those attributes — model provider, model name, inference type, inference mode, pricing unit — are standardized and available by default. That’s the plumbing work no one talks about, but it’s what makes downstream reporting actually reliable,” Chopra added. This, said Jain, makes it easier to build dashboards showing cost by model, provider, token type or inference mode while also identifying expensive workloads, unusual token growth and opportunities to move to cheaper models or batch processing. It’s a timely update, especially in light of last week’s AWS billing issue that caused some customers to see incorrect cost estimates of services consumed in the AWS Management Console, said Muskan Bandta , cloud associate at FinOps services providing firm ZopDev. “Anything that gives customers clearer, more granular and more trustworthy billing data is welcome when confidence in the numbers has just been shaken. It does not fix what went wrong, but better visibility into where spend is going is exactly what teams want more of after an episode like that,” Bandta added. This article first appeared on InfoWorld .
- The latest Chinese AI models may indeed work for enterprises, but only in a handful of specific applications
Ever since Chinese AI startup DeepSeek launched three years ago, enterprise executives have been nervous about relying on Chinese AI models . But now that the latest Chinese AI offerings, Alibaba’s 2.4-trillion-parameter model Qwen3.8 Max and Moonshot’s 2.8-trillion-parameter model Kimi K3 , are promising even more powerful performance, those IT executives are being forced to again ask if these models are worth using, even in a limited fashion. Former Walmart risk official Steven Eric Fisher , now an independent cybersecurity and risk advisor, thinks they should at least take another look. “Enterprises should take these models seriously, but neither adopt nor reject them solely because they are Chinese,” he said. “They should be assessed like any other critical technology dependency: jurisdiction, ownership, training and software provenance, licensing, data handling, hosting, security, reliability, and the ability to independently test their behavior. Geopolitical exposure is a legitimate risk factor, but it should be incorporated into technical and supply-chain diligence rather than used as a substitute for it.” Choose applications with care He added, “Chinese models may be especially valuable for coding, multilingual processing, high-volume document analysis, research, synthetic-data generation, and privately operated security or forensic workflows, but they should be subject to task-specific testing rather than broad benchmark claims.” Shashi Bellamkonda , principal research director at Info-Tech Research Group, agreed that the Chinese models can work well if they are only used in carefully chosen applications. “Although Moonshot’s K3 still trails Claude’s Fable 5 and GPT 5.6 Sol on performance and user experience, good companies that have governance and prompt guardrails will not face the instability and improvisation of [the Chinese] models,” he said. “These models will win in usage. US frontier models are leading as the best models, but Chinese models will be sufficient for high-volume, low-drama tasks that cost less for non-critical transactions.” On the flipside, Bellamkonda suggested a variety of areas where enterprises should avoid Chinese AI models, including “customer-facing work without a human in the loop, regulated or sensitive data, and anything where a hallucinated answer creates legal or safety exposure. That is where the reliability gap and the political-radioactivity concern both bite, and where the closed American models still earn their premium.” Bellamkonda said he didn’t see the differences in data reliability, mostly involving hallucination rates, as meaningful for enterprise AI strategy decisions. “Every open-weight model in this class can get facts wrong or make things up. That is fixable with the right setup, so it is not a reason to avoid these models,” he said. “For high-volume tasks with clear limits, you feed the model your own trusted documents to answer from, and you keep a person checking the output. That combination is safe for production. The model on its own is not.” Too early for enterprises to consider However, not everyone agrees that the latest Chinese models have earned their place as enterprise AI decision options. Cybersecurity consultant Brian Levine , executive director of FormerGov, focused on Chinese technology concerns when he worked for the US Justice Department as its representative in the US law enforcement Joint Liaison Group (JLG) with China. “It is way too early for US enterprises to seriously consider these models,” he said. “Until proven otherwise, enterprises should assume that if they use these models, they may be granting China complete access to everything they do through the models, and potentially access to their networks and employees more broadly. At this point, any pros of using such models are strongly outweighed by the potential security, confidentiality, and reliability concerns.” Tom Findling , CEO of Conifers.ai, was equally emphatic that enterprise CIOs need to steer clear of these newer Chinese models. “Using them inhouse? Absolutely not. You simply don’t know what is planted inside of it and you don’t know what training data is put into them,” Findling said. Mike Wilkes , enterprise CISO at Aikido Security, added that the very attractive pricing for these Chinese models may be appealing, but suggested that, despite the low cost, they’re ultimately too risky. “Enterprises should take these models seriously, but not romantically. Parameter count is horsepower measured in a showroom, not braking distance in the rain,” he said. “The real tests are reliability on your data, the cost of a wrong answer, and whether the model behaves predictably under pressure.” He noted that the benchmarks on the latest open-weights models are impressive, and very close to those of the frontier lab models, which makes the cost ”incredibly seductive, especially when a team does not want to risk their data being used to train those frontier models.” But the Chinese models can still work in specific circumstances. “The strongest value will be in bounded, reversible and inspectable work: coding inside a sandbox, multilingual translation, document triage, data extraction and other high-volume tasks where outputs can be verified,” he said. “Cheap intelligence is valuable, but only when it is not mistaken for trustworthy judgment.” Wilkes added that the regulatory issues surrounding Chinese models can be especially problematic. Texas, for example, has banned their usage . A rational choice for some workloads However, Yuri Goryunov , CIO of consulting firm Acceligence, argued that CIOs should seriously consider these models. “Counterintuitively, the biggest benefit of Kimi and models like it is the lack of guardrails,” Goryunov said. “Think of it as stick shift cars in the era of automatics. If you want ease and comfort, stay with the frontiers because they have cruise control, shift the gears for you and they decide when. If you want performance and control, expand your horizons. But a stick shift assumes you know how to drive one: you bring your own governance, your own evals, your own safety layer. That’s a cost and specialized talent, which is super rare, and for the right organization it’s also the whole point.” Goryunov’s bottom line: “For internal, high-volume, well-harnessed workloads, [the Chinese models] have moved from ‘watch list’ to ‘rational choice.’”
- The EU’s AI transparency deadline is weeks away. Is your enterprise ready?
Providers and deployers of AI systems: You only have a couple of weeks left until you must explicitly inform users when they are interacting with AI content. To assist in the effort, the European Commission (Commission) has published guidelines to help AI deployers get in line with the AI Act’s transparency obligations, which will begin to go into effect on August 2. After that, companies providing AI systems must alert users when they are interacting with AI. They must also tell users when they have been exposed to deepfakes, “emotion recognition,” or biometric categorization systems, or when they are given AI-manipulated content in matters of “public interests without human review or editorial control.” Henna Virkkunen , the Commission’s executive VP for tech sovereignty, security and democracy, said in a statement, “with today’s guidelines, the Commission supports the smooth and effective application of the AI Act to make AI systems interacting with people such as chatbots and AI agents and AI content more transparent and trustworthy. These guidelines support providers and deployers in meeting their obligations under the AI Act, while helping citizens know when they are interacting with AI.” Systems must include machine-readable markers to reveal such content, to reduce “the risk of deception and manipulation” and build public trust in AI. “Generative systems have collapsed the cost of producing convincing content while the cost of judging it stands where it always stood,” said Sanchit Vir Gogia , chief analyst at Greyhound Research. This requirement is “an attempt to restore friction to that imbalance.” A company’s non-compliance could result in fines anywhere from €750K (about $856K) to €15M (about $17 million), or even up to 3% of its total worldwide annual revenue. Transparency requirements The EU AI Act’s transparency requirements apply to “natural or legal persons,” public authorities, agencies, or other bodies that develop AI systems, or have them developed, and place them on the EU market or into use under their name or trademark. This means all companies, regardless of whether or not they are EU-based. “Systems placed on the European market, put into service there, or producing outputs used there are inside the field, wherever the developer sits,” Gogia noted. Applicable systems must be intended to interact directly with “natural persons”; these systems include AI-enabled chatbots or conversational agents, AI companions, or coding agents. However, AI-enabled tools like recommender systems, spam filters, authentication, search and retrieval, transcription, text and code auto-completion, or predictive maintenance do not fall under the rule. Specific outputs such as AI-generated text, images, video, and audio must contain a machine-readable mark. Deepfakes and public interest-related text created by AI without human review or control must be clearly labeled, however, deepfake content that is “artistic, creative, satirical, or fictional” is largely exempt. AI content must be marked with one of three labels: “AI,” “Fully AI-generated,” or “Partially AI-modified.” For instance, “Fully AI-generated” applies when news summaries, music, art, or videos have been created without any human oversight (apart from prompting), while “partially AI-modified” could mean a person’s face is swapped into an authentic photograph to create a deepfake. The three icons are publicly available for free use; enterprises can download zip files in PNG and SVG formats. Most of the Act’s transparency rules begin to go into effect on August 2. But AI systems placed on the market before then will have some leeway; they must be in compliance by December 2. However, a four-month allowance “on one obligation, for one population of systems, contingent on one procedural step, is not a strategy,” Gogia emphasized. Enterprises should plan to comply by August 2 and “treat any relief that arrives as margin.” A consistent code of practice Along with the transparency guidelines, the Commission has introduced a code of practice that essentially serves as a gesture of good faith. When signed, it can provide “legal certainty” and a “simple and practical” way to demonstrate compliance with the AI Act , according to the Commission. Signatories can also collaborate through the ‘Signatory Taskforce,’ which will share practices and advance technologies around marking and labeling practices. Providers that choose not to sign must comply through other methods and demonstrate that those methods are “adequate” through assessment by surveillance authorities, according to the Commission. Non-signatories “keep their flexibility, and will face more case-by-case scrutiny for it,” said Gogia. Criteria for compliance Shashi Bellamkonda , principal research director at Info-Tech Research Group, pointed out that the transparency requirements apply to content only when three criteria are met: It has been published, is informative to the public, or is on matters of public interest. B2B business content or blogs may not need an AI disclosure if they do not meet these criteria, he noted. Also, published text that has undergone human review or is under editorial control does not need to be labeled. Editorial control means that a person must hold the ultimate legal responsibility for the publication of the content. Many companies like Google, Adobe, and LinkedIn have already established ways to identify images marked as AI-generated. Meta has made it a requirement, but the creator has to add the AI-generated label, Bellamkonda said. “This is a good move for guardrails around public information, and companies with good compliance and ethical oversight may not have to worry about this,” he noted. But as a general practice, companies should disclose AI-generated content and state whether it has been human reviewed. Creating a transparency pipeline Establishing full transparency means identifying who carries the responsibility for the content, whether the marking survives real use, not just testing, and what evidence will defend the decision, Gogia said. Concerns cluster around responsibility, durability and evidence. Several organizations usually touch one piece of content, and none controls the whole chain, which is why contracts become the “pressure point,” he said. Most current agreements were written to deliver software and say “almost nothing” about provenance persistence, verification access, or evidence retention. The durability concern is the most difficult, Gogia noted, because marking performs well in controlled settings but “badly in ordinary life.” Meta, for one, said its invisible watermark was designed to survive cropping; a published test, however, found the company’s preview detector missed 55% of cropped images . “CIOs should ask which platform can actually provide evidence before believing its dashboard,” said Gogia. Disclosure of AI use must be “clear, distinguishable and accessible,” he emphasized. “A notice buried in lengthy terms, or reachable only through determined clicking, satisfies nobody, least of all a market surveillance authority.” Sustained compliance is a “living control” requiring a central record of systems, duties and evidence; testing taking place where the user meets the control rather than where the developer built it; and continuous supplier assurance. Enforcement will vary by country, so keep one common baseline with local overlays, Gogia said. His advice: Inventory every system that talks to people, generates content, or gauges sentiment; classify provider and deployer roles; place disclosures at first interaction; define substantive human review; keep the evidence. Marks and provenance signals should be tested after content undergoes cropping, compression, translation, transcription, and other editing, Gogia said. A useful audit starts from a real output and follows its “pulse” through generation, editing and publication, identifying at “each beat” the responsible party, the surviving mark, and evidence for exceptions. Missed labels should also be traced for root cause and recurrence. To ensure compliance, before August 2, enterprises need a prioritized inventory, live disclosures on the highest-risk use cases, and a “named owner for every control,” he noted. In the first 30 days, they should stabilize and test; in the first 90 days, push requirements into procurement processes as a standing discipline. Procurement must secure commitments on marking methods, known failure modes, and evidence access, with explicit notice if/when any of them change. “The sensible architecture is a common transparency baseline carrying traceability, responsibility, and evidence, with jurisdictional overlays for language, sector rules, and local practice,” Gogia said. This article originally appeared on CIO.com .
- Build a voice agent with LiveKit and AssemblyAI’s Voice Agent API
Tutorial on integrating LiveKit with AssemblyAI’s Voice Agent API to build a voice agent.
- AI Is Rewriting Cybersecurity's Rules
AI Is Rewriting Cybersecurity's Rules Anonymous (not verified) Mon, 07/20/2026 - 20:00 Dateline Tue, 07/21/2026 - 12:00 Mercury ID 691203 Summary Sentence AI is giving cybercriminals new speed and scale, but researchers say the technology could also become one of cybersecurity’s strongest defenses. Story Link Learn More Core Research Areas Artificial Intelligence at Georgia Tech Cybersecurity
- Responsible AI Is Becoming a Growth Strategy
A playbook for turning good governance into a competitive edge.
- Ecosystem Roundup: SEA’s AI future is being coded in its own languages
Southeast Asia‘s AI moment is arriving, but not in English. Across Vietnam, Thailand, and Indonesia, a quiet but consequential shift is under way: researchers, startups, and governments are building large language models trained on local languages, dialects, and cultural contexts, rather than waiting for Silicon Valley to localise its tools. The stakes are significant. English-centric […] The post Ecosystem Roundup: SEA’s AI future is being coded in its own languages appeared first on e27 .
- YouTube reveals which AI slop videos cant make money
YouTube clarified its AI slop rules and outlined three types of low-quality or risky videos that cannot earn money.