AI News Archive: July 21, 2026 — Part 16
Sourced from 500+ daily AI sources, scored by relevance.
- Nvidia Pitches DLSS 5 Again, Promises to Preserve 'Artistic Intent'
Nvidia Pitches DLSS 5 Again, Promises to Preserve 'Artistic Intent' PCMag UK
- How to Watch Samsung's London Unpacked Keynote: New Foldables and Snapdragon AI Watches Expected
How to Watch Samsung's London Unpacked Keynote: New Foldables and Snapdragon AI Watches Expected PCMag Australia
- Microsoft and Mistral strike multi-billion-dollar deal to build AI infrastructure across Europe
Microsoft and Mistral are expanding their strategic partnership with a multi-billion-dollar deal to build out AI infrastructure in Europe. The article Microsoft and Mistral strike multi-billion-dollar deal to build AI infrastructure across Europe appeared first on The Decoder .
- Microsoft Strikes Deal For Mistral's AI Computing Power
Microsoft Strikes Deal For Mistral's AI Computing Power Barron's
- London robotics startup Humanoid raises $152M Series A, with Bosch set to manufacture its wheeled robots at scale
London-based robotics startup Humanoid has raised $152 million in a Series A round led by Prime Movers Lab, the venture firm behind Figure AI’s $39 billion valuation, pushing Humanoid’s own valuation above $1 billion. Bosch, Schaeffler, Aglaé Ventures, the investment arm of LVMH chairman Bernard Arnault, and Taiwan’s Fubon Financial also participated. The round brings […] This story continues at The Next Web
- UK-based Humanoid secures $152m Series A funding
The financing is the largest-ever Series A round for a humanoid-first robotics company in Europe, the company said. Read more: UK-based Humanoid secures $152m Series A funding
- China's Moonshot AI is seeking a $50 billion valuation in pre-IPO funding talks
The Chinese AI startup plans to open discussions in August on a final fundraising round before a Hong Kong listing as soon as this year
- Iceland's Sowilo secures pre-seed backing to expand AI fashion product platform Catecut - ArcticStartup
Iceland's Sowilo secures pre-seed backing to expand AI fashion product platform Catecut - ArcticStartup ArcticStartup
- Even Healthcare in talks to raise $50 million from Khosla Ventures, Alpha Wave
The round will likely be at a valuation of about $300 million, nearly double that of $153 million it saw in its last funding round in January. The company decline to comment on the fundraise.
- Traveltech Startup 30 Sundays Raises ₹61 Cr To Expand AI-Powered Holiday Planning
Traveltech startup 30 Sundays has raised ₹61 Cr (about $6.7 Mn) in a Series A round led by Bessemer Venture…
- Google Releases Three New Gemini A.I. Models
The models include one that is the company’s most powerful and another that is fine-tuned for cybersecurity, as Google competes with rivals like OpenAI and Anthropic.
- Google releases three new Gemini models — but no 3.5 Pro
Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber, but the continued absence of Gemini 3.5 Pro raises fresh questions about its AI strategy.
- Google announces Gemini 3.6 Flash and cybersecurity AI, teases 3.5 Pro and Gemini 4
There are new 3.6 and 3.5 models today, but Google is already training Gemini 4.
- Google's Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks —and 3.5 Pro is on the way
Google DeepMind today released three new proprietary AI models it says are among its most token-efficient yet: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The models aim to make AI agents faster, smarter, and cheaper at scale. Google is pricing Gemini 3.6 Flash at $1.50 per one million input tokens and $7.50 per one million output tokens through its application programming interface (API), while Gemini 3.5 Flash-Lite costs a staggeringly cheap $0.30/$2.50 per million tokens in/out. Compare that to the $1.50/$9.00 per 1M tokens for Gemini 3.5 Flash, and the $2/$12 for Gemini 3.1 Pro Preview, and the savings are considerable. However, Google's prior generation Gemini 3.1 Flash-Lite still remains the search giant's "most cost-efficient" model at $0.25/$1.50 per 1M tokens. Yet, it remains 2X slower than the new, more expensive Gemini 3.5 Flash-Lite, giving those enterprises who value speed more "bang" for their buck. VB Frontier AI Model API Pricing Comparison Chart (Late July 2026 Shortlist) Model Input ($/1M) Output ($/1M) Total ($/1M) Source MiMo-V2.5 Flash $0.10 $0.30 $0.40 Xiaomi deepseek-v4-flash $0.14 $0.28 $0.42 DeepSeek deepseek-v4-pro $0.435 $0.87 $1.305 DeepSeek MiniMax-M3 $0.30 $1.20 $1.50 MiniMax LongCat-2.0 — limited-time promo $0.30 $1.20 $1.50 LongCat Gemini 3.1 Flash-Lite $0.25 $1.50 $1.75 Google Qwen3.7-Plus $0.40 $1.60 $2.00 Alibaba Cloud MiMo-V2.5 $0.40 $2.00 $2.40 Xiaomi Gemini 3.5 Flash-Lite $0.30 $2.50 $2.80 Google LongCat-2.0 — standard $0.75 $2.95 $3.70 LongCat MiMo-V2.5 Pro (≤256K) $1.00 $3.00 $4.00 Xiaomi GLM-5.2 $1.40 $4.40 $5.80 Z.ai GPT-5.6 Luna $1.00 $6.00 $7.00 OpenAI Grok 4.5 $2.00 $6.00 $8.00 xAI MiMo-V2.5 Pro (>256K) $2.00 $6.00 $8.00 Xiaomi Gemini 3.6 Flash $1.50 $7.50 $9.00 Google Qwen3.7-Max $2.50 $7.50 $10.00 Alibaba Cloud Gemini 3.5 Flash $1.50 $9.00 $10.50 Google Gemini 3.1 Pro Preview (≤200K) $2.00 $12.00 $14.00 Google GPT-5.6 Terra $2.50 $15.00 $17.50 OpenAI GPT-5.4 $2.50 $15.00 $17.50 OpenAI Kimi K3 $3.00 $15.00 $18.00 Moonshot AI Gemini 3.1 Pro Preview (>200K) $4.00 $18.00 $22.00 Google Claude Opus 4.8 $5.00 $25.00 $30.00 Anthropic GPT-5.5 $5.00 $30.00 $35.00 OpenAI GPT-5.5 Instant (chat-latest) $5.00 $30.00 $35.00 OpenAI Sakana Fugu Ultra (≤272K) $5.00 $30.00 $35.00 Sakana AI GPT-5.6 Sol $5.00 $30.00 $35.00 OpenAI Claude Fable 5 / Claude Mythos 5 $10.00 $50.00 $60.00 Anthropic No price was provided yet for the specialty Gemini 3.5 Flash Cyber model, which, as its name would imply, is designed for cybersecurity researchers and red teamers to patch bugs. While the prices are among the middle-low end of all major AI models globally, the fact that Google designed them to use less tokens overall also should drive down costs for enterprises beyond what the sticker price shows (since you'll be paying for fewer total tokens at any rate). Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available immediately through the Gemini API in Google AI Studio and Android Studio, as well as within the consumer Gemini application and Google Search. According to a separate Google blog post , Gemini 3.5 Flash Cyber will be available "exclusively available to governments and trusted partners via CodeMender soon" — CodeMender being Google's proprietary AI code bug-fixing agent released last year. As with previous Gemini models, these are all proprietary and "closed source," thus, they can only be obtained through Google's official API and that of its partners, as opposed to an open-source license like MIT or Apache 2.0. One conspicuous omission noted by developers on X and social media: where is the larger, more powerful, flagship Gemini 3.5 Pro model Google previously alluded would be released this summer? After all, Gemini 3.1 Pro, the prior flagship, debuted back in February 2026 , and rivals OpenAI and Anthropic have since released several more generations of flagship updates far more powerful than Google's. Google technical staffer Logan Kilpatrick responded to one such inquiry on X, writing : "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it’s ready." Google's release signals that the immediate future of AI lies in agentic capabilities—systems that operate autonomously over extended periods. If early large language models are akin to massive, fuel-hungry freight trains capable of hauling incredible loads at immense cost, the new Flash series represents a fleet of nimble, hyper-efficient hybrid delivery vans. Efficiency gains ranging from 17% to 65% reduced tokens for strong results on third-party benchmarks Under the hood, Gemini 3.6 Flash achieves significant efficiency gains. The model reduces output token usage by 17% compared to its predecessor, Gemini 3.5 Flash, according to the Artificial Analysis Index maintained by the independent third-party AI benchmarking group of the same name. In specific long-horizon software engineering benchmarks like DeepSWE , which measures how well agents complete multi-step engineering tasks from scratch, the token savings reach up to 65%. This reduction means the model requires fewer reasoning steps and tool calls to complete the exact same multi-step workflow. Think of token efficiency like fuel economy in a vehicle. When an AI model takes a convoluted path to solve a problem, it burns through more computational fuel, driving up the final cost for the developer. By streamlining its internal logic, Gemini 3.6 Flash arrives at the correct answer faster and cheaper. While Google's materials did not specify the exact architectural or algorithmic changes used to achieve this token efficiency, they noted that the model "takes fewer reasoning steps and tool calls to accomplish multi-step workflows" and exhibits reduced "verbosity." The official model cards released by Google reveal that both Gemini 3.6 Flash and Gemini 3.5 Flash-Lite feature a 1-million-token input context window alongside a max output limit of 64,000 tokens, with both models sharing a knowledge cutoff date of March 2026. Respectable benchmark performance at low cost The technological improvements extend to concrete capabilities. Gemini 3.6 Flash scores 49% on the DeepSWE benchmark, a notable increase from the 37% achieved by version 3.5. It also pushes machine learning engineering performance higher, scoring 63.9% on MLE-Bench compared to 49.7% previously. Furthermore, Google integrates computer use as a built-in client-side tool via the Gemini API and Gemini Enterprise, reflecting an OSWorld-Verified score of 83.0%, up from 78.4%. The model also tackles knowledge work with greater proficiency, outperforming its predecessor on benchmarks like GDPval-AA v2 by moving from a score of 1349 to 1421. To ensure safety amidst these capability upgrades, Google deploys enhanced Frontier Safety safeguards. These protections harden the model against jailbreaks and mitigate risks in Chemical, Biological, Radiological, and Nuclear domains, as well as cyber offense misuses. The engineering team trains the model to minimize refusals for beneficial uses, striking a necessary balance between strict security and practical utility. M odels for low-cost coding, agentic, and cybersecurity use cases — respectively Google divided its new offerings into three distinct products tailored for different operational needs. Gemini 3.6 Flash serves as the heavy-duty workhorse of the trio. It handles complex coding, intricate knowledge work, and multimodal processing with improved precision. Enterprise customers utilize it for demanding tasks such as complex document parsing, intricate chart and data analysis, and long-form report drafting. The model executes complex code migrations using multi-agent orchestration frameworks with lower latency and higher quality than earlier iterations. Furthermore, 3.6 Flash aids in developing photographic texture extractors for 3D workflows using canvas interfaces. Gemini 3.5 Flash-Lite targets environments where high throughput and absolute minimal latency are non-negotiable. Google designates it as the fastest model in the 3.5 series. As measured by Artificial Analysis, the model processes 350 output tokens per second, making it highly effective for agentic search and massive document processing workloads. Artificial Analysis notes this is about twice as fast as prior generation model Gemini 3.1 Flash-Lite. Developers can configure 3.5 Flash-Lite to prioritize low-latency execution for high-volume tasks using minimal thinking levels, or engage higher thinking levels to process complex multi-step subagent workloads. Despite its lite designation, it outperforms the standard Gemini 3 Flash on several key agentic and coding evaluations, including SWE-Bench Pro, where it scores 54.2% compared to 49.6%, and OSWorld-Verified, scoring 74.0% versus 65.1%. The model extracts product features from massive datasets, generates interactive web design concepts, and scales receipt translation seamlessly. The third product, Gemini 3.5 Flash Cyber, represents a highly specialized deployment. Google fine-tuned this model specifically to find and fix cybersecurity vulnerabilities. It integrates directly with Google's CodeMender agent. In practice, multiple 3.5 Flash Cyber agents work concurrently to produce a single, comprehensive vulnerability report, achieving competitive performance at the frontier on the CyberGym benchmark, even getting within range of Anthropic's much-hyped Mythos model. Google did not specify an exact numerical cost for 3.5 Flash Cyber, stating only that it is fine-tuned "at a lower price per token than larger models. Commercial licensing only The licensing framework for the new Gemini models carries profound implications for developers and enterprise users. Google deploys Gemini 3.6 Flash and 3.5 Flash-Lite under a commercial, proprietary API model. Unlike open-source software governed by licenses such as the MIT License or the GNU General Public License, developers do not gain access to the underlying model weights, training data, or source code. An MIT or GPL license grants users the freedom to download the codebase, modify the internal architecture, self-host the deployment, and distribute the software infrastructure independently. In contrast, Google's API approach means developers essentially rent access to the intelligence on a strict metered basis. Every prompt and generated response travels through Google's managed servers, incurring a cost based on the strict pricing structure of $1.50 per million input tokens for 3.6 Flash. This commercial tethering restricts deployment flexibility. Enterprises cannot air-gap the models entirely on their own local secure hardware without establishing specialized, high-tier enterprise agreements with Google Cloud. Developers remain bound by Google's acceptable use policies, arbitrary rate limits, and network requirements, creating a permanent dependency on Google's infrastructure uptime and terms of service. The licensing for Gemini 3.5 Flash Cyber proves even more restrictive. Acknowledging the dual-use nature of cybersecurity AI—which attackers can weaponize just as easily as defenders can use it to patch systems—Google is for now making the model only available behind a limited-access pilot program, similar to the trend kicked off by Anthropic's Mythos model with its Project Glasswing program , and continued by OpenAI with its staggered rollout for GPT-5.6 . In this case, Google is making 3.5 Flash Cyber exclusively available to governments and trusted partners. This strict gatekeeping prevents open access, prioritizing systemic security over widespread developer innovation. Looking ahead Google DeepMind continues to iterate rapidly, but the gap in its product line remains apparent. While the Flash series excels in speed and economy, the industry eagerly awaits the deployment of Gemini 3.5 Pro to gauge Google's absolute frontier capabilities. Simultaneously, the company confirms that pre-training for Gemini 4 has already commenced. Until the next major flagship release materializes, developers must optimize their systems using the highly efficient, yet purposefully constrained, Flash architecture.
- Google Releases New Gemini Flash Models, But Flagship Still Delayed
Google Releases New Gemini Flash Models, But Flagship Still Delayed The Information
- Google is launching three new AI models, including a security tool to rival Anthropic
Gemini 3.5 Flash Cyber, built to find and patch software vulnerabilities, will initially be available only to governments and trusted partners
- Google updates lightweight Gemini models, but flagship still delayed
Google launched three new Gemini AI models, offering cost-effective solutions. The company has not provided a release date for its flagship Gemini 3.5 Pro. This top-tier model's delay is reportedly due to internal performance shortcomings. Google's rivals have recently introduced their own advanced artificial intelligence models. The company stated Gemini 3.5 Pro is currently undergoing partner testing.
- Google ships three new Gemini Flash models but its frontier 3.5 Pro remains lost in training
Google is shipping three new Flash models in the Gemini series, including the more efficient 3.6 Flash, which uses up to 65 percent fewer tokens, and a cybersecurity model available only to governments and select partners. But the anticipated flagship, Gemini 3.5 Pro, is still missing, while OpenAI, Anthropic, and Chinese labs are already competing at the frontier level. The article Google ships three new Gemini Flash models but its frontier 3.5 Pro remains lost in training appeared first on The Decoder .
- Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads
Google Releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A Cheaper, More Token-Efficient Flash Tier Built for Agentic Workloads MarkTechPost
- Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
- Gemini 3.6 Flash and Gemini 3.5 Flash-Lite: Halving Time per Task
Google’s Gemini 3.6 Flash and 3.5 Flash‑Lite cut task completion time by half.
- Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on AI Gateway
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on AI Gateway. Gemini 3.6 Flash improves quality across coding, agentic tasks, and web development while consuming fewer tokens and making fewer model calls. It produces cleaner web and app development output. Gemini 3.5 Flash Lite upgrades the agentic capabilities of the Flash-Lite tier, making it a good fit for subagents that handle scoped parts of a larger task. To use them, set model to google/gemini-3.6-flash or google/gemini-3.5-flash-lite in the AI SDK : AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting , Zero Data Retention support , budgets for API keys , routing rules , and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Gemini 3.6 Flash in the model playground . Read more
- Google Introduces Gemini 3.6 to Remind You It Has an AI Model, Too
Plus: 3.5 Flash-Lite, and 3.5 Flash Cyber.
- Google Releases 3 New Gemini Models, 3.5 Pro Still Not Available
The new models prioritize token efficiency and performance.
- Google expands Gemini 3.5 line with trio of new models — and shares an update on Gemini 3.5 Pro
Meet Google's latest Gemini models.
- Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4
As we wait for 3.5 Pro, Google today announced Gemini 3.6 Flash and 3.5 Flash-Lite, while providing updates on what comes next.
- Google launches a cheaper alternative to large AI security models like Mythos
Google is launching Gemini 3.6 Flash alongside a new security model dedicated to quickly finding and patching security vulnerabilities. In a blog post on Tuesday, Google describes Gemini 3.5 Flash Cyber as a "cost-efficient and highly capable alternative" to larger, more expensive AI systems, such as the one offered by Anthropic's Mythos. The cybersecurity model […]
🤖 ModelsJul 21, 2026https://www.theverge.com/tech/968572/google-gemini-flash-cyber-ai-security-model - Google’s Gemini 3.6 Flash targets enterprise agent token costs
Google has released Gemini 3.6 Flash and 3.5 Flash-Lite as new workhorses designed to cut latency and token costs for enterprise AI agents. The economics of running autonomous software agents inside a production environment come down to a fixed equation few vendors advertise directly. A model needs to reason through a multi-step task competently, but […] The post Google’s Gemini 3.6 Flash targets enterprise agent token costs appeared first on AI News .
- New Gemini 3.5 Flash Models Are Faster and Cheaper but Not Smarter
The updated models are intended to be more affordable for enterprises. The new cyber model is designed to orchestrate.
- Google releases two new Gemini models, but still no Gemini 3.5 Pro
On July 21, Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and a new cybersecurity model, while Gemini 3.5 Pro remains unreleased.
- Poolside releases Laguna S 2.1, the open-weight coding model pitched as the West’s answer to DeepSeek and Qwen
Poolside has released Laguna S 2.1, a 118-billion-parameter open-weight model built for agentic coding that the San Francisco startup says matches or exceeds models several times its size. The model uses a mixture-of-experts architecture with eight billion active parameters per token, is compact enough to run on a single Nvidia DGX Spark desktop system, and […] This story continues at The Next Web
- Laguna S 2.1 is now available on AI Gateway
Laguna S 2.1 from Poolside is now available on AI Gateway. There are 2 versions of the model available: Free version (256K context window): poolside/laguna-s-2.1-free Paid version (1M context window): poolside/laguna-s-2.1 Laguna S 2.1 is an open-weight Mixture-of-Experts model that supports a context window of up to 1M tokens and runs in thinking and no-thinking modes. The model specializes in agentic coding and long-running tasks, including writing and debugging code, running tests, building browser-based tooling, and working on MLOps pipelines and AI research. In thinking mode, Laguna S 2.1 reports 70.2% on Terminal-Bench 2.1, 78.5% on SWE-bench Multilingual, and 59.4% on SWE-Bench Pro. To use Laguna S 2.1, set model to poolside/laguna-s-2.1-free or poolside/laguna-s-2.1 in the AI SDK : AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting , Zero Data Retention support , budgets for API keys , routing rules , and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Laguna S 2.1 in the model playground . Read more
- Kimi K3 model highlights China's burgeoning AI efforts
China’s AI were on full display over the last week.
- Another ‘DeepSeek moment’? What China’s Kimi K3 means for the global AI industry
The launch of Moonshot AI’s Kimi K3 has revived an intense debate that has raged in Silicon Valley ever since DeepSeek’s shock breakthrough last year: whether China can overcome its limited access to advanced chips to match the performance of the United States’ frontier artificial intelligence models. Trillions of US dollars might rest on the answer, as some see Chinese developers’ success in increasing performance through architectural innovation as weakening the rationale for America’s vast...
- What China's Internet Is Saying About Moonshot's Hot New Kimi Model
What China's Internet Is Saying About Moonshot's Hot New Kimi Model Business Insider
- Moonshot AI K3 Triggers AI Trading Reckoning: Winners and Losers in the New Landscape Reshaped by 2.8 Trillion Parameters
Kimi K3 launch triggers global AI stock revaluation: Chinese semiconductor and memory chip makers win, OpenAI and Anthropic valuations lose $314B, and model developers face brutal competition.
- China Just Released Two AI Models That Claim to Rival OpenAI and Anthropic. And They're Free.
China Just Released Two AI Models That Claim to Rival OpenAI and Anthropic. And They're Free. entrepreneur.com
- Cisco releases Antares, open-weight small models for locating code vulnerabilities
Cisco Systems Inc. today introduced Antares, a family of small language models built to pinpoint where known security vulnerabilities sit inside a codebase, and released the first two as open-weight downloads on Hugging Face. The models come from Cisco Foundation AI, the company’s research and engineering group focused on security-specific artificial intelligence. Antares targets vulnerability […] The post Cisco releases Antares, open-weight small models for locating code vulnerabilities appeared first on SiliconANGLE .
- Alibaba Qwen 3.8 Max Shows China Closing in on U.S. Models
The low-cost, open-weight model and others from China give enterprises more choices, given the performance claims of some Chinese model providers.
- The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation
Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Models (LLMs), including downstream tasks, such as Neural Machine Translation (NMT). With the latest research indicating that RLVR could be the preferred training m...
- Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, existing studies rely primarily on prompt engineering or supervised fine-tuning, while systematic research on reinforcement learning (RL) post-training and automat...
- MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument ...
- Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning
Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and continuous parameterization. Prior work established the effectiveness of single-agent actor-critic algorithms - Greedy Actor-Critic (GAC), Soft Actor-Critic (SAC), an...
- SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation
Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question. Existing datasets and benchmarks are limited in scale, domain coverage, or executable verification, leaving the true gap between current LLMs and reliable scienti...
- On the Effectiveness of Pretraining for Graph Combinatorial Optimization
This paper introduces a self-supervised pretraining framework for graph combinatorial optimization specifically designed to address the nature of routing problems like the Traveling Salesman Problem. By utilizing graph contrastive learning with geometric augmentations (specifically, rotations and ax...
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a l...
- Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs
Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistics lacks explanatory power regarding the source of er...
- Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions
Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work shows that six moderation classifiers achieve at most 20.9 to 24.5% accuracy and nine state-of-the-art VLMs remain at or below 10.2% with illusion-aware prompting, l...
- Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts
Fast prediction of the response of adhesive soft viscoelastic contacts represents a current challenge in soft robotics and for gripping and manipulation tasks. Determining the complete time-resolved force trajectory requires full numerical simulations, whose computational cost is strongly parameter-...
- Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE language model, AdamW keeps 50.6 GB of first and second moments to update 12.6 GB of bfloat16 weights. We study SkewAdam, an optimizer built on the observation that the...