AI News Archive: August 29, 2026 — Part 2
Sourced from 500+ daily AI sources, scored by relevance.
- Caught in 4K: The Aurora Files
Executive Summary An exposed open directory revealed months of activity from belonging to a Russian speaking Aurora ransomware affiliate, active against more than twenty organisations between April and July 2026. The directory included the ... (https://incidentdatabase.ai/cite/1661#7855)
Score: 45🌐 MovesAug 29, 2026https://www.cloudsek.com/blog/aurora-ransomware-affiliate-ai-attack-planning-crypto-payments - Glow-Up or AI Slop? Nvidia's Controversial DLSS 5 Leaks Early Through Mods
Glow-Up or AI Slop? Nvidia's Controversial DLSS 5 Leaks Early Through Mods PCMag
Score: 45🌐 MovesAug 29, 2026https://www.pcmag.com/news/glow-up-or-ai-slop-nvidias-controversial-dlss-5-leaks-early-through-mods - Warnock blasts OpenAI, Effingham for doing data center deal ‘in the dark’
Warnock blasts OpenAI, Effingham for doing data center deal ‘in the dark’ AJC.com
- Marvell selloff deepens as investors seek clarity on Google AI deal payoff
Marvell has become a market darling fueled by the AI spending boom as Big Tech races to adopt custom chips for greater cost efficiency and performance, powering its shares to nearly triple this year.
- How AI‑powered policing is endangering the public's trust in police
Since the release of ChatGPT in 2022, artificial intelligence (AI) has dominated the cultural landscape. Uptake across various sectors, including law enforcement, quickly followed, usually with minimal foresight and little to no guardrails, like effective policies or laws, to govern AI use.
Score: 42🌐 MovesAug 29, 2026https://phys.org/news/2026-08-aipowered-policing-endangering-police.html - Bill Gates in his own words: How he’s using AI, and why he’s worried about the future
Bill Gates warned this week that the AI industry is crossing safety lines it set for itself, and that no one is preparing for what comes next. On this episode of the GeekWire Podcast, we play highlights from our interview with Gates about the essay, and debate his three proposals. Read More
Score: 42🌐 MovesAug 29, 2026https://www.geekwire.com/2026/bill-gates-in-his-own-words-how-hes-using-ai-and-why-hes-worried-about-the-future/ - The Intricacies Of Artificial General Intelligence And Curing Disease
Artificial general intelligence (AGI) may enable human-level reasoning and critical thinking across multiple cognitive domains.
- ICE plans to spend up to $2M on robot dogs
The four-legged robots are equipped with cameras and Wi-Fi connectivity and can travel at speeds of up to about 5 feet per second
Score: 41🌐 MovesAug 29, 2026https://www.the-independent.com/news/world/americas/us-politics/ice-robot-dogs-trump-immigration-b3041608.html - There is a growing opposition to data centres globally: Shylesh Muralidharan, Portulans Institute
There is a growing opposition to data centres globally: Shylesh Muralidharan, Portulans Institute
- Global investors use “perp” bets to chase China tech stocks like Unitree
Risky derivatives offer overseas traders a way to tap the hottest IPOs.
Score: 39🌐 MovesAug 29, 2026https://kr-asia.com/global-investors-use-perp-bets-to-chase-china-tech-stocks-like-unitree - Meet the humanoid robots cleaning homes for $30 an hour
These $30-an-hour humanoid robots are cleaning homes in San Francisco, tackling tasks like mopping floors, tidying rooms and cleaning kitchens. NBC News' Tom Llamas speaks with Tau Robotics CEO and co-founder Alexander Koch about how the robots work and the challenges of making them fully A.I. powered.
Score: 38🌐 MovesAug 29, 2026https://www.nbcnews.com/video/humanoid-robots-cleaning-homes-for-30-an-hour-269041733517 - AI is breaking the find-and-fix model for application security
Contrast Security recently released AppSec Overflow 2026, a research report showing that the find-and-fix workflow underpinning modern application security no longer holds up against AI-accelerated attackers and AI-powered security assessments. […] The post AI is breaking the find-and-fix model for application security appeared first on Express Computer .
Score: 38🌐 MovesAug 29, 2026https://www.expresscomputer.in/news/ai-is-breaking-the-find-and-fix-model-for-application-security/138169/ - Anthropic is cutting Claude Code's current weekly limits by 17%
Anthropic is permanently increasing Claude Code's standard weekly usage limits by 25% for Pro, Max, Team, and seat-based Enterprise plans, but it's not as good as it sounds. [...]
- A week with the Google Pixel 11 changed my mind about Gemini (for the better)
In a year of iterative upgrades, Google is introducing smart features that make the base Pixel the standout in its generation.
- Apple CarPlay gaining a new AI chatbot app before Siri AI arrives
Since iOS 26.4, Apple has supported conversational apps on CarPlay. This category includes AI chatbot apps. Siri AI joins the CarPlay party starting with iOS 27. The iPhone software update is currently in beta ahead of the official launch in September. Meanwhile, the fourth third-party AI chatbot has arrived with Apple CarPlay support: Meta AI.
Score: 35🌐 MovesAug 29, 2026https://9to5mac.com/2026/08/29/apple-carplay-gaining-a-new-ai-chatbot-app-before-siri-ai-arrives/ - Siri AI Can Only Do So Much. For Smart Home Dominance, Apple Needs to Improve These Two Products, and Fast
Siri AI Can Only Do So Much. For Smart Home Dominance, Apple Needs to Improve These Two Products, and Fast PCMag Australia
- Why Ambient Clinical Scribes Drop Pertinent Negatives: Architecting Dual-Pass Extraction Control…
Why Ambient Clinical Scribes Drop Pertinent Negatives: Architecting Dual-Pass Extraction Control Towers for Medical AI How middle-context attention degradation in long conversational transcripts causes silent omission errors in EHR SOAP notes, and how to build deterministic audio timeline reconciliation gates. Ambient clinical scribes represent one of the fastest-growing enterprise deployments of generative AI in healthcare. By capturing doctor-patient conversational audio, transcribing dialogue, and structuring clinical interactions into SOAP (Subjective, Objective, Assessment, Plan) notes, these systems offer immense administrative time savings. However, moving from conversational summarization to mission-critical clinical documentation exposes a fundamental structural failure in large language models: the Silent Omission Vulnerability . Consider a 25-minute outpatient cardiology encounter. The patient presents with intermittent palpitations. During the review of systems, the physician asks: “Have you noticed any shortness of breath, chest pressure, or swelling in your lower legs?” The patient responds: “None at all. No chest pain, no shortness of breath, and no swelling.” The ambient AI scribe processes the audio transcript, generates a well-formatted SOAP note, and stages it for physician signature. Yet, when the note is reviewed, the Review of Systems (ROS) section details the palpitations and blood pressure readings, but completely omits the patient’s explicit denial of chest pain and shortness of breath. In clinical medicine, a pertinent negative is an essential diagnostic boundary that rules out acute coronary syndrome or heart failure. Omitting this data creates severe clinical blind spots, compromises longitudinal patient care, and creates substantial malpractice exposure. +--------------------------------------------------------------------------------------------------+ | THE SILENT OMISSION FAILURE IN SINGLE-PASS SCRIBES | +--------------------------------------------------------------------------------------------------+ [Clinical Audio Encounter] ──► 25 Minutes of Doctor-Patient Dialogue (4,500 Tokens) │ ▼ [Single-Pass LLM Scribe] ──► Single Prompt: "Extract entities and format as SOAP note" │ ├────────────────────────────────────────────────┐ ▼ ▼ [Chief Complaint & Plan] [High Attention: Tokens 0-500 & 4000-4500] [Captured & Synthesized] │ ▼ [Intermediate Negatives] ──► [ATTENTION SINK: Tokens 1500-3000] ──► [DROPPED / OMITTED SILENTLY] ("Denies chest pain/SOB") │ ▼ [Generated SOAP Note] ──► Grammatically Fluent Note with Empty Review of Systems │ ▼ [Clinical Consequence] ──► Lost Diagnostic Baseline / Malpractice Risk 1. The Core Architectural Failure Vectors in Ambient Scribes Evaluating generative AI in clinical documentation reveals that over 80% of severe documentation errors are not hallucinations, but errors of omission . This failure stems from three architectural vectors: Failure Vector A: Middle-Context Attention Sinks in Long Transcripts Conversational clinical transcripts are lengthy, noisy, and unstructured. A standard 20-minute consultation produces between 3,000 and 6,000 tokens of raw dialogue. In transformer-based architectures, self-attention mechanisms exhibit the well-documented “lost in the middle” phenomenon: models allocate significantly higher attention weights to tokens at the absolute start and end of the context window, while attention degrades sharply across the intermediate tokens. Intermediate review of systems dialogue is frequently bypassed during token generation. Failure Vector B: Semantic Bias Against Negative Assertions Foundational models pre-trained on generic internet corpora are optimized for information density. In everyday conversation, negative statements often represent filler or non-events. However, in medical ontology (such as SNOMED-CT or ICD-10), an explicit negation (pertinent_negative) has identical diagnostic weight to an active finding (pertinent_positive). Single-pass LLMs routinely compress or discard negated phrases during narrative synthesis because they treat negative statements as absence of data. Failure Vector C: Cognitive Automation Bias in Clinical Sign-Off Because generative scribes produce grammatically coherent, professional notes, clinicians experience automation bias. When reviewing a clean draft at the end of a long clinical shift, physicians scan for false additions (hallucinations) far more effectively than they spot missing data (omissions), resulting in unverified commits to the live Electronic Health Record (EHR). 2. The Deterministic Governance Architecture To eliminate omission errors, ambient documentation architectures must abandon single-pass generation in favor of a Dual-Pass Extraction Engine coupled with Deterministic Timeline Reconciliation . STATEFUL DUAL-PASS AMBIENT SCRIBE CONTROL TOWER [Raw Clinical Audio Encounter & Word-Level Timestamp Transcript] │ ├───────────────────────────────────────────────┐ ▼ ▼ ┌──────────────────────────────────────────────────┐ ┌──────────────────────────────────────────────────┐ │ PASS 1: ENTITY EXTRACTION ENGINE │ │ PASS 2: NARRATIVE SYNTHESIS ENGINE │ │ • Dedicated Extraction Prompt │ │ • Structural SOAP Note Formatting │ │ • Strict Pertinent Positive/Negative Schema │ │ • Clinical Style & Tone Harmonization │ │ • Binds Each Entity to Audio Timestamp Spans │ │ • Generates Draft Subjective/Objective Text │ └─────────────────────────┬────────────────────────┘ └─────────────────────────┬────────────────────────┘ │ │ └───────────────────────┬─────────────────────────────┘ │ ▼ ┌────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ DETERMINISTIC TIMELINE RECONCILIATION GATE │ │ • Cross-references extracted Negative/Positive entities against synthesized SOAP narrative │ │ • Verifies that all confirmed transcript entities exist in final clinical draft │ └─────────────────────────────────────────────────┬──────────────────────────────────────────────────────┘ │ [All Entities Reconciled?] / \ YES/ \NO ▼ ▼ ┌────────────────────────────────────────────────────┐ ┌────────────────────────────────────────────────┐ │ STAGED EHR CLINICAL DRAFT NOTE │ │ EXECUTION CIRCUIT BREAKER │ │ • Structured Note Ready for Clinician Review │ │ • Halt Automated Staging │ │ • Inline Audio Provenance Tooltips Attached │ │ • Highlight Missing Negatives in UI Diff │ │ • Atomic One-Click Commit to Production FHIR EHR │ │ • Alert Clinician to Verify Omitted Findings │ └────────────────────────────────────────────────────┘ └────────────────────────────────────────────────┘ 3. Production Implementation: Dual-Pass Extraction & Reconciliation Engine The following Python implementation demonstrates how an enterprise control tower decouples entity extraction from synthesis and deterministically reconciles clinical findings against transcript coordinates: from pydantic import BaseModel, Field, ConfigDict from typing import List, Optional, Set from enum import Enum import logging logging.basicConfig(level=logging.INFO) logger = logging.getLogger("AmbientScribeControlTower") class AssertionType(str, Enum): POSITIVE = "POSITIVE" NEGATIVE = "NEGATIVE" class ClinicalEntity(BaseModel): model_config = ConfigDict(extra="forbid", frozen=True) concept_name: str = Field(..., min_length=2) snomed_code: Optional[str] = Field(None, pattern=r"^\d{6,18}$") assertion: AssertionType audio_start_sec: float = Field(..., ge=0.0) audio_end_sec: float = Field(..., ge=0.0) transcript_segment: str = Field(..., min_length=2) class ExtractedClinicalState(BaseModel): model_config = ConfigDict(extra="forbid", frozen=True) encounter_id: str = Field(..., min_length=3) entities: List[ClinicalEntity] class SynthesizedSOAPNote(BaseModel): model_config = ConfigDict(extra="forbid", frozen=True) encounter_id: str = Field(..., min_length=3) subjective: str objective: str assessment_and_plan: str class ScribeReconciliationGateway: def reconcile_and_stage_note( self, extracted_state: ExtractedClinicalState, synthesized_note: SynthesizedSOAPNote ) -> dict: """ Deterministically verifies that all extracted clinical entities (especially negatives) are explicitly represented in the synthesized SOAP note prior to clinician staging. """ combined_note_text = ( f"{synthesized_note.subjective} {synthesized_note.objective} " f"{synthesized_note.assessment_and_plan}" ).lower() omissions: List[ClinicalEntity] = [] for entity in extracted_state.entities: # Deterministic check: verify concept presence in synthesized note if not self._is_entity_represented(entity, combined_note_text): omissions.append(entity) # Circuit breaker trigger if pertinent findings were dropped if omissions: self._trip_omission_circuit_breaker(extracted_state.encounter_id, omissions) return { "status": "CIRCUIT_BREAKER_TRIPPED", "encounter_id": extracted_state.encounter_id, "omission_count": len(omissions), "dropped_entities": [e.model_dump() for e in omissions] } logger.info( f"SUCCESS: Encounter {extracted_state.encounter_id} fully reconciled with zero omissions." ) return { "status": "STAGED_FOR_SIGNATURE", "encounter_id": extracted_state.encounter_id, "synthesized_note": synthesized_note.model_dump() } def _is_entity_represented(self, entity: ClinicalEntity, note_text: str) -> bool: """ Validates whether an extracted entity and its negative/positive assertion are present in the note body. """ concept = entity.concept_name.lower() if concept not in note_text: return False # If negative, verify that negation syntax co-occurs in the text if entity.assertion == AssertionType.NEGATIVE: negation_terms = ["no ", "denies", "denied", "negative for", "without", "free of"] return any(term in note_text for term in negation_terms) return True def _trip_omission_circuit_breaker(self, encounter_id: str, omissions: List[ClinicalEntity]) -> None: logger.error(f"CRITICAL: Omission Circuit Breaker Tripped for Encounter {encounter_id}.") for item in omissions: logger.error( f"DROPPED {item.assertion.value}: '{item.concept_name}' " f"spoken between {item.audio_start_sec}s - {item.audio_end_sec}s " f"in transcript: \"{item.transcript_segment}\"" ) # In production: Flag note in EHR UI diff view and alert attending physician 4. Architectural Comparison: Single-Pass Scribing vs. Stateful Control Towers Moving Beyond Single-Pass Ambient Documentation Generative AI provides extraordinary efficiency gains in clinical documentation, but single-pass models cannot guarantee completeness across complex, multi-turn clinical encounters. Relying solely on physician review to catch missing data invites clinical errors due to automation bias. Governing high-stakes healthcare AI requires rigorous infrastructure: separating entity extraction from narrative synthesis, binding all extracted concepts to audio timestamps, and deterministically reconciling final notes before they ever reach an EHR. Published by Maya Lin On the team at Claire By The Algorithm Why Ambient Clinical Scribes Drop Pertinent Negatives: Architecting Dual-Pass Extraction Control… was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- Musicians-turned-detectives are hunting for AI grifters
As audio-focused generative tools and platforms have gotten more sophisticated, the internet has become increasingly filled with AI-generated music whose melodies and vocals are algorithmically derived from the work of human artists. While some of the people pumping out this kind of content immediately own up to using AI, others have denied using the technology […]
Score: 35🌐 MovesAug 29, 2026https://www.theverge.com/entertainment/985866/h4rris-nihil-young-edm-suno-ai - An AI Oracle’s Rise and Fall
Plus, homeowners who are mad about mold, a contemplative camino in Spain, and why employers should lower their expectations in November.
Score: 34🌐 MovesAug 29, 2026https://www.wsj.com/tech/ai/an-ai-oracles-rise-and-fall-9b0cebea?mod=rss_Technology - How to use the new Siri app in iOS 27
How to use the new Siri app in iOS 27.
- AI can make China’s political language sound more threatening, scholar warns
China is being urged to use human translators to refine AI renderings of its political messaging, amid concerns that literal translations could make Beijing’s language sound more threatening abroad. An article in a state-affiliated academic newspaper said the country should strengthen its translation capacity and guard against AI-generated distortions to improve how its policies and intentions were understood overseas. As artificial intelligence (AI) takes on more translation work, human...
- Gemini Spark is almost a dream AI assistant — except for 1 thing Perplexity does better
Using Perplexity Computer extensively revealed exactly how limited Gemini Spark is.
Score: 32🌐 MovesAug 29, 2026https://www.androidauthority.com/gemini-spark-vs-perplexity-computer-3703364/ - Why logistics AI is becoming the bext big bet for investors
Why logistics AI is becoming the bext big bet for investors YourStory.com
Score: 32🌐 MovesAug 29, 2026https://yourstory.com/2026/08/why-logistics-ai-is-becoming-the-bext-big-bet-for-investors- - FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning irep.mbzuai.ac.ae
Score: 32🌐 MovesAug 29, 2026https://irep.mbzuai.ac.ae/bitstreams/0c859596-27d1-4679-a792-9ad51d83975b/download - Charts that explain the rise of artificial intelligence
Models keep improving, tech companies keep climbing on the stock market, and hundreds of millions of people are using a technology that only recently arrived
Score: 31🌐 MovesAug 29, 2026https://english.elpais.com/technology/2026-08-29/charts-that-explain-the-rise-of-artificial-intelligence.html - Startup bags $7M to build drone-interceptor-in-a-backpack systems
Spike drone defense system is designed to be carried in a rucksack or mounted on top of a vehicle
- Truckers are finally making real money again — and AI is a big reason why
Truckers are finally making real money again — and AI is a big reason why
Score: 30🌐 MovesAug 29, 2026https://www.marketwatch.com/bulletins/redirect/go?g=bfc73fb6-f5f9-4d31-a745-232b26f05219&mod=mw_rss_bulletins - These cyborg cockroaches could be a lifesaver for people trapped in the rubble following natural disasters
Australian researchers have equipped cockroaches with miniature rescue technology, demonstrating controlled movement and emergency delivery capabilities inside difficult environments.
- Gen Z's AI Edge: Tech leaders hail digital natives’ workplace value; HP and LinkedIn back skills
HP executive George Brasher refutes claims of Gen Z laziness and interpersonal skill deficits. He highlights their digital native status as a significant workplace advantage. This generation possesses strong AI skills, enhancing job performance and work-life balance. LinkedIn co-founder Reid Hoffman and Stripe's Emily Glassberg Sands also acknowledge Gen Z's tech adaptability. Their AI fluency positions them well for entry-level roles in the evolving job market.
- Genrobotics signs MoU with Territorial Army for defence, disaster-response robotics
The lineup also included Genbot, designed for critical missions requiring human-level accessibility in infrastructure, a robotic rover and G-KAI, a rehabilitation device for upper-body recovery
- How AI Is Changing Car Accident Claims – What Every Driver Should Know
How AI Is Changing Car Accident Claims – What Every Driver Should Know USA Today
- Former Lyft Drivers Now Doing Menial Tasks for Their Waymo Overlords
They're along for the ride. The post Former Lyft Drivers Now Doing Menial Tasks for Their Waymo Overlords appeared first on Futurism .
Score: 28🌐 MovesAug 29, 2026https://futurism.com/advanced-transport/former-lyft-drivers-menial-tasks-waymo - ‘They find themselves obsessed, forgoing sleep and self-care’ — what ‘AI psychosis’ looks like, and why experts question the term
More and more people are experiencing delusional thinking fuelled by AI, but how should we talk about it?
- Beyond Critical Thinking: The Verification Process AI Actually Needs
Skeptical intelligence helps AI users be more productive with AI. It is an improvement on critical thinking.
- You no longer need a subscription to try Google’s Dreambeans app
Once exclusive to Google AI Ultra subscribers, Dreambeans is now free for any US Google Account on Android and iOS.
Score: 27🌐 MovesAug 29, 2026https://www.digitaltrends.com/computing/you-no-longer-need-a-subscription-to-try-googles-dreambeans-app/ - Apple @ Work: Parallels Desktop 27 brings OpenGL 4.3 and AI acceleration to Apple Silicon
Apple @ Work is exclusively brought to you by Mosyle , the only Apple Unified Platform. Mosyle is the only solution that integrates in a single professional grade platform all the solutions necessary to seamlessly and automatically deploy, manage, and protect Apple devices at work. Over 45,000 organizations trust Mosyle to make millions of Apple devices work ready with no effort and at an affordable cost. Request your EXTENDED TRIAL today and understand why Mosyle is everything you need to work with Apple . Virtualization on macOS has become rock-solid in the Apple Silicon era. Parallels Desktop 27 pushes that further with a new Metal-based graphics driver, OpenGL 4.3 support for Windows virtual machines on Apple silicon, faster local AI workloads, and a host of enterprise features aimed at IT teams, helping them manage and deploy Windows on macOS. Here’s my look at what’s new after spending some time with it.
- Farther Data Centers Can Deliver Faster AI Responses as Networks Get Smarter
Farther Data Centers Can Deliver Faster AI Responses as Networks Get Smarter USA Today
- ClearWay Mobility Introduces Obstacle Detection Device to be Added to the Standard White Cane Used by Visually Impaired People
ClearWay Mobility Introduces Obstacle Detection Device to be Added to the Standard White Cane Used by Visually Impaired People
- ChatGPT, Gemini, and Claude quietly get worse over time. This hidden setting fixes them
Your favorite AI chatbot might be remembering a little too much.
Score: 24🌐 MovesAug 29, 2026https://www.androidauthority.com/chatgpt-gemini-claude-memory-cleanup-how-3701167/ - AI e-commerce projects may fail without clean, centralised data
E-commerce businesses need to focus on developing solid data frameworks prior to investing in artificial intelligence. The success of AI ventures is closely tied to data readiness, and poor data quality can lead to substantial challenges, hampering AI effectiveness. Companies should strive for centralized data architectures to create tailored customer experiences. By addressing data issues upfront, firms can expedite AI rollouts and reduce the risk of project failures.
- VerSe Innovation unveils global AI content platform SparkStation for studios, brands and creators
VerSe Innovation, the parent company of Dailyhunt and Josh, has launched SparkStation, an AI-powered content production platform that brings different stages of video and advertising production into a single workflow. The company unveiled the platform at Film Expo 2026 as it looks to target filmmakers, production houses, brands, agencies, creators and influencers. SparkStation covers the production process from ideation and scripting to storyboarding, casting, location and costume design, shot creation, music and dialogue generation, editing, captioning, rendering and ad creation. The platform is designed to allow users to move between these stages without switching between multiple tools. For filmmakers, SparkStation includes tools such as story builder, script coverage, casting, location design and shot creation. Users can also generate music and dialogue and edit and render the final output. Importantly, projects can be exported to professional editing software including Premiere Pro, Final Cut Pro and DaVinci Resolve, allowing production teams to retain their existing post-production workflows. The platform also addresses consistency, which remains a challenge with AI-generated video. Users can lock five elements, including a character’s face, expression, wardrobe, location and props, and carry them across multiple scenes. Spark Station uses a model-agnostic AI layer to assign different tasks to models suited to them. Video generation, character creation, camera movements, voice and language can therefore be handled by different models while remaining within the same interface. For brands, e-commerce companies and agencies, the platform supports image, video, UGC and performance ad creation. It can generate product creatives at catalogue scale and produce multiple versions of campaigns for different platforms, languages and formats. SparkStation also supports localisation in more than 60 languages, including changes to lip movements, emotional delivery and local expressions. For creators and influencers, the company is offering templates and studio-style production tools through a credits-based model. Studios and brands will have access to seat-based SaaS plans, while larger customers can opt for enterprise licensing and API access. VerSe claims SparkStation can reduce the cost of producing a 60-second brand film from Rs 10-25 lakh to Rs 50,000-75,000. It also expects production timelines to fall from several weeks to around 24 hours. These figures are company estimates and will need to be tested as the platform scales. SparkStation enters beta today, with general availability from October 1. VerSe will invest over $30 million in 24 months, targeting over 100 studios, 1,000 brands and thousands of creators.
Score: 24🌐 MovesAug 29, 2026https://entrackr.com/news/verse-launches-sparkstation-to-cut-ai-content-production-costs-by-90-12451424 - When to Use Claude Code and When to Use Codex
Learn where which coding agent is best The post When to Use Claude Code and When to Use Codex appeared first on Towards Data Science .
Score: 23🌐 MovesAug 29, 2026https://towardsdatascience.com/when-to-use-claude-code-and-when-to-use-codex/ - How to Run a Chatbot on Your Own Computer
Installing a large language model on your personal computer gives you a handy digital assistant that won’t compromise your data privacy.
- This pocket AI voice recorder took me back to the good old days of dictation (and shorthand)
The Comulytic Note Pro AI voice recorder packs a lot of power into a credit card-sized device. But its usefulness is up to you.
Score: 22🌐 MovesAug 29, 2026https://www.zdnet.com/article/comulytic-note-pro-ai-voice-recorder-review/ - SFT, RL and DPO: The Other Stack
Post-training is the part that produces the weights everything else in this series then has to serve. Different tooling, different failure modes, and one connection to the serving stack that decides how long a training run takes. A bonus chapter, not part eight. Everything below the arrow is the rest of this series. Part 1 : Start Here: The Words Everyone Uses About LLM Inference Part 2 : Your KV Cache Is Bigger Than Your Model Part 3 : What a Kernel Is, and Why Everyone Is Writing New Ones Part 4 : Quantization Is Four Decisions, Not One Part 5 : Tuning vLLM: What Every Setting Does to the Arithmetic Part 6 : Your Second GPU Is Bought for the Cache, Not the Model Part 7 : Your Context Length Decides What a Kernel Is Worth This series has been about serving a model. What it costs, where the time goes, which settings matter. All of it assumed the model already existed. Somebody handed you a set of weights and your job started there. Post-training is where those weights got their behaviour. SFT, DPO, PPO, GRPO and RLVR are the methods, and it is a different job from serving, done with different tools and usually by different people. Most of the time post-training and serving never touch, which is fine. Reinforcement learning is the exception. It makes its own training data by running the model. That means a training run is full of inference, and the speed of your serving stack decides how long the training takes. Where post-training sits Pre-training is the enormous, expensive part: read most of the internet, learn to predict the next token. What comes out completes text and is close to useless as an assistant. Post-training turns that into something you would ship. It is comparatively cheap, and almost all the behaviour you care about comes from here. When people talk about fine-tuning their own model, this is usually what they mean. Then serving, which is the other seven parts. Five boxes left to right: most of the internet, pre-training, a model that completes text, post-training, something you would ship. Pre-training stops at the middle box. Post-training picks it up from there. SFT: show it what good looks like Supervised fine-tuning is the plain one. Collect examples of the behaviour you want, a prompt and the response you wish the model had given, and train on them directly. Same objective as pre-training, but on curated data. It teaches format, instruction-following, tone and domain register. It is cheap, well understood, and still where most of the practical value is . LoRA (low-rank adaptation) is how most people afford it. It freezes the base weights and trains a small pair of matrices alongside them, so what you update is a fraction of the model. The saving is not really the parameter count. A frozen weight carries no gradient and no optimizer state, and that is where training memory goes. LoRA also changes what you walk away with. A full fine-tune hands you a new model. LoRA hands you an adapter, a small file that sits beside the base weights, and one server can hold several of them at once and pick one per request. So post-training and serving meet here too, not just at the rollouts later in this chapter. Most people will hit this one rather than the rollouts, because it arrives with the method they are already using. None of this is specific to SFT. The same trick works under DPO and under the reinforcement learning methods, and it changes how many networks you have to hold. DPO needs a frozen copy of the model you started from. PPO and GRPO need a frozen reference. Under LoRA you already have one, because the base weights never moved, so switching the adapter off gives you the reference for nothing. The variants mostly differ in how the adapter itself is parameterised, which matters far less than the decision to use one at all. QLoRA is the exception worth knowing. It holds the frozen base at four bits and trains the adapter on top. That is part four's quantization turning up in the training stack, and it often decides whether a large model fits on hardware you own. SFT trains on an answer you wrote yourself. The small block beside the model is the LoRA adapter. When a model is badly behaved and the plan is reinforcement learning, the SFT data is the cheaper place to look first. Reinforcement learning cannot install a format or a tone the examples never showed. DPO: learn from comparisons Eventually you want the model to prefer one answer over another where you cannot write the ideal response, only recognize it. That needs comparisons: this answer is better than that one. The old way was reinforcement learning from human feedback: train a reward model on the comparisons, then use RL to push the model toward high reward. It works, and it is a lot of machinery. Direct preference optimization removes most of it. You feed it triples: a prompt, the better answer, the worse one. It never builds the reward model and never runs the RL loop. Instead it optimizes one closed-form objective that reaches the same optimum the long route would, provided preferences follow the Bradley-Terry model. Same target, reached offline. That is not a promise of the same result on the data you actually have. In practice it is a classification loss against a frozen copy of the model you started with. DPO measures the trained model against a frozen copy of where it started, and that comparison is what keeps it from drifting. That is why it became the default for preference work: no reward model to train, no rollouts to generate, and one loss function to tune. PPO and GRPO: when you actually need reinforcement learning Comparison cannot teach everything. Sometimes what you want is not a style but a correct answer reached by a chain of reasoning. That is what reinforcement learning is for. PPO , proximal policy optimization, is the classical approach. Four networks are resident: the policy being trained, a reward model to score what it produces, a frozen reference to keep it from drifting, and a critic . The critic predicts how well a state should turn out, so the algorithm can tell whether an outcome beat expectations or fell short of them. There are four networks here, and only two of them are trained. The policy and the critic carry gradients and optimizer state, which is where the memory goes. The reward model and the reference only ever run forward, so you can offload them or put them in another process. GRPO is the one that drops the critic. A critic roughly doubles what you have to train. GRPO removes the critic. Instead of learning to predict the baseline, it samples a group of answers to the same prompt, scores them all, and uses the group's own mean as the baseline. Answers above it get reinforced, answers below it discouraged. The advantage is the score minus the group mean, divided by the spread. Numbers here are examples, but the subtraction is what GRPO actually does. RLVR makes this practical for reasoning. It stands for reinforcement learning with verifiable rewards. That is a different thing from reinforcement learning from human feedback, where people do the scoring. Here a program does it, so if the task has a checkable answer you need no reward model at all. A maths answer key, a unit test suite, a compiler. The reward is 0 or 1 from a program rather than a learned score, so it cannot drift the way a learned reward model does. However, it can still be rigged: models special-case unit tests, exploit answer-extraction formats, and reach right answers by invalid routes the checker never sees. But then the problem is a hole in your checker, which is a program you can open and read. You cannot read a reward model the same way. GRPO with a verifier on top is the published recipe behind DeepSeek-R1 and the open reasoning models that followed it. Which one, and why it is usually not a close call The rows are in the order these methods were invented. Which one you use depends on what data you can get. Work down the list: Start with SFT. One network, no sampling, and the only method that works when you can write the answer down. If the problem is format, tone or domain knowledge, nothing further down this list will fix it better. Add DPO when you can rank but not write. A second network in memory, still no sampling, covering the subjective tasks that have no correct answer. Cheaper than anything below it on this list. Reach for GRPO with a verifier when correctness is checkable. Sampling is where the cost jumps, because generating G answers per prompt is a serving workload. It is not the only route to a reasoning gain. Plain SFT on reasoning traces from a stronger model has produced large gains on its own. There is also a live argument that RLVR sharpens what the base model already samples rather than adding capability. Its advantage over distillation is that it needs no stronger teacher. Consider PPO last, and probably not at all. Four resident networks against GRPO’s two. That two assumes the reward comes from a verifier. GRPO removes the critic, not the reward model, so with a learned reward model it is three. Either way you are paying for an advantage that only shows up when credit has to be assigned token by token across a long sequence. If you can neither check an answer nor rank it against another, reinforcement learning has no signal to work with at all. Your problem is the SFT data. The connection to everything else in this series Reinforcement learning rollouts are inference. A GRPO run samples a group of answers per prompt, eight or sixteen or thirty-two, for every prompt in every batch, for the whole run. That generation is ordinary decoding. It uses the model and the KV cache the way serving does, and it runs into the memory bandwidth limit from part one on the engines part five described. Most of your post-training time goes there. That means the arithmetic in this series applies to a training run. The KV cache still decides how big a batch you can hold, and rollouts have an advantage here: you generate G answers for every prompt, so the batch runs into the hundreds. Part one put the ridge on an H100 at 296 and said almost nobody sees it in practice. A rollout batch clears it, on a dense model at least. On a mixture of experts each expert only sees its own share, so the number that matters is smaller than the headline. The answers in a group also start from the same prompt, so they share a prefix exactly. Part two called this often the biggest single win available on the workloads that have one. Quantization is where the comparison stops. Part two’s doubling came from halving the cache . Halving the weights is a different thing. It hands back a fixed number of bytes once, and what that buys depends on how your budget is split. On part two’s numbers, 65.25 GB of weights become 32.6, freeing about 30 GiB, so a 71 GiB cache budget becomes roughly 101. That is 1.4×, not 2×. Two caveats. Those are serving numbers, borrowed to show the size of the effect rather than because a 120B policy and its optimizer state would fit on two cards. And gpt-oss-120b already ships its experts at four bits, so for this checkpoint the halving is hypothetical. Quantizing a rollout policy also costs something that quantizing a served model does not. The rollout policy is the model being trained. Quantize it and the version that generated the samples is no longer quite the version being updated. The objective works by comparing those two, so the comparison stops being trustworthy. In serving, quantization costs you quality, which is what part four is about and which you settle with your own evaluation. In a rollout you pay that and the broken comparison on top. Weights and cache share one budget here, so halving the weights returns less than halving the cache did in part two. The KV budget is tighter than part two’s as well, because in a colocated run the cache only gets what is left after weights, gradients, optimizer state and activations. Training and serving hit the same limit. What to do with this Fix the SFT data before reaching for anything with “reinforcement” in the name. That is cheaper, faster, and sometimes, more effective. It is also usually where the real problem was. If you are running rollouts at any scale, profile the generation separately from the training step. If most of your time is in generation, and it very often is, the seven preceding parts of this series are about your training job too. Conclusion Post-training and serving get treated as two disciplines with two vocabularies. SFT and DPO have nothing to do with paged attention or tensor parallelism. But the moment a method generates its own samples to learn from, it is doing inference, and everything the rest of this series says about inference applies to it. So it is worth knowing whether yours does. All five methods drawn side by side: SFT, DPO, PPO, GRPO and RLVR, each showing where its training signal comes from. The whole chapter on one image. What decides the method is your use case and what you have to check the work with. If you have run any of these, I would like to hear which one and how it went. Please clap or comment and follow me. The rest of this series is about serving the model rather than training it, and it starts with part one . SFT, RL and DPO: The Other Stack was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
Score: 22🌐 MovesAug 29, 2026https://pub.towardsai.net/sft-rl-and-dpo-the-other-stack-0ab7026d528e?source=rss----98111c9905da---4 - Superblog Opens an API and MCP Server So AI Agents Can Publish Directly to a Blog
Superblog Opens an API and MCP Server So AI Agents Can Publish Directly to a Blog USA Today
- The Pixel 11 is a swing-and-a-miss of a phone, but this one AI camera feature is genuinely useful
The Pixel 11 is a swing-and-a-miss of a phone, but this one AI camera feature is genuinely useful Tom's Guide
- Visa Expands Support for its Clients and the Industry as Organisations Navigate New AI Era of Cybersecurity
Visa Expands Support for its Clients and the Industry as Organisations Navigate New AI Era of Cybersecurity
- RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need
Enterprise Document Intelligence [Vol.1 #B00] - Retrieval answers one kind of question. Classifying a request, matching free text to a reference list, reading a table, cleaning OCR noise: each has a cheaper method that works, and the engineering is knowing which one to reach for The post RAG Is Not the Whole Toolkit: The NLP Techniques Real Problems Still Need appeared first on Towards Data Science .
Score: 20🌐 MovesAug 29, 2026https://towardsdatascience.com/rag-is-not-the-whole-toolkit-the-nlp-techniques-real-problems-still-need/ - What's the difference between TPU vs. GPU?
Google's Pixel 11 phone uses a Tensor G6 processor with a powerful TPU. How is it different from a GPU, and what does that mean in real-world use?
Score: 20🌐 MovesAug 29, 2026https://www.engadget.com/2242759/tpu-vs-gpu-difference-between-processsors/