AI News Archive: May 21, 2026 — Part 17
Sourced from 500+ daily AI sources, scored by relevance.
- AI data centers are booming. Why Texas may be best prepared
AI data centers are booming. Why Texas may be best prepared Austin American-Statesman
- AI-powered traffic cameras debut near Austin
AI-powered traffic cameras debut near Austin Austin American-Statesman
- Your employees’ hidden AI habits that put your business at risk
Your employees’ hidden AI habits that put your business at risk Raleigh News & Observer
- Microsoft storms RAMPART, adds Clarity to agentic AI safety
Redmond open sources two tools for building and maintaining safer agents
- Researchers develop AI-powered stretchable computing patch
Researchers develop AI-powered stretchable computing patch EurekAlert!
- AI not yet good enough to grade university essays, rewarding ‘style over substance’
AI not yet good enough to grade university essays, rewarding ‘style over substance’ EurekAlert!
- Widespread generative AI use demands reform in higher education assessment
Widespread generative AI use demands reform in higher education assessment EurekAlert!
- Widespread AI misuse by college students signals need to rethink assessment
Widespread AI misuse by college students signals need to rethink assessment EurekAlert!
- Revisiting the detection, fate, and health risks of microplastics in the environment through artificial intelligence
Revisiting the detection, fate, and health risks of microplastics in the environment through artificial intelligence EurekAlert!
- SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data
We propose SADGE, a quantitative similarity metric that predicts the performance of synthetic image datasets for common computer vision tasks without downstream model training. Estimating whether a synthetic dataset will lead to a model that performs well on real-world data remains a bottleneck in m...
- MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often yields unnatural or implausible outcomes, especially by missing secondary causal consequences. To address this, we intro...
- MiddAI
Offline AI with online search & memory capability.
- DocInsightHub AI
Compliance Tracking with AI
- Private AI Assistant
Private AI Assistant
- AionDB
Database, Graph, Vector, Rag, AI, SQL
- Ateegnas
An AI Native Data Annotation Tool
- BetterVideo.io
AI video enhancement. Credits never expire. Data privacy
- Maya by Authority AI
Your next investor is already in your network
- Disene
Turn any word into your own understanding
- Seleqt
AI Qualify leads and personalize outreach across channels
- Dialfyne
AI that answers your leads and trains your sales team
- Careerbot
A career assistant in plain markdown.
- Zenso
Your brand, your products, in every AI answer.
- infobro.ai
Expert AI Tool Reviews, News & Insights
- Tokenwatch
Codex/Claude Code local token watcher
- Resurface — Your web, remembered.
Find anything you've seen online using natural language
- LawCaseAI
AI legal case management for modern lawyers
- Swapd AI
AI-powered fashion rental marketplace for Gen Z
- TrustBoost PII Sanitizer
Context-aware PII sanitization for autonomous AI agents
- Selfcoder - Local AI coding assistant
Local AI coding assistant powered by LM Studio or Ollama.
- RogueProof
The all-in-one social proof engine for modern brands
- ClientScrape AI Pro
Autonomous AI B2B Lead Generation & Bulk Email Extractor
- AILatest Journal
Journal rankings and indexing signals in one place
- WhatsApp Marketing Solution
AI-driven messaging solutions
- The Futureproof Copywriter Audit
AI-proof your freelance writing career in 60 minutes
- Luuzon
Secure AI CRM automating tenant screening & fraud detection
- xMap POI Data
POI data built for AI agents, apps, & location intelligence.
- InstaVM
Instant computers for AI agents
- - G42
G42
- Tokenisation via Convex Relaxations
Tokenisation is an integral part of the current NLP pipeline. Current tokenisation algorithms such as BPE and Unigram are greedy algorithms -- they make locally optimal decisions without considering the resulting vocabulary as a whole. We instead formulate tokeniser construction as a linear program ...
- Vector Policy Optimization: Training for Diversity Improves Test-Time Search
Language models must now generalize out of the box to novel environments and work inside inference-scaling search procedures, such as AlphaEvolve, that select rollouts with a variety of task-specific reward functions. Unfortunately, the standard paradigm of LLM post-training optimizes a pre-specifie...
- Reducing Political Manipulation with Consistency Training
Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from opposing political sides asymmetrically. We refer to this phenomenon as covert political bias and identify 7 categories of techniques through which ...
- Understanding Data Temporality Impact on Large Language Models Pre-training
Large language models (LLMs) are typically trained on shuffled corpora, yielding models whose knowledge is frozen at train time and whose temporal grounding remains poorly understood. In this work, we study the impact of pre-training dynamics on the acquisition of time-sensitive factual knowledge, f...
- Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationalised by the TRUST multi-agent large language model (LLM) pipeline. Using a Bundestag plenary speech by Felix Banaszak (51 segments, 245 s) as a case st...
- AMEL: Accumulated Message Effects on LLM Judgments
Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing through one conversation. We ask whether the polarity of prior conversation history biases subsequent judgments, an effect we call the accumulated messa...
- Self-Policy Distillation via Capability-Selective Subspace Projection
Self-distillation bootstraps large language models (LLMs) by training on their own generations. However, existing methods either rely on external signals to curate self-generated outputs (e.g., correctness filtering, execution feedback, and reward search), which are costly and unavailable for the be...
- Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
Previous detection studies have shown that LLMs cannot be effectively used as detectors, but these studies have not addressed modern Chinese poetry. Moreover, no relevant research has explored the performance of LLMs in detecting modern Chinese poetry. This paper evaluates and enhances the performan...
- Multi-Stage Training for Abusive Comment Detection in Indic Languages
In recent years social media has become an increasingly popular tool for communication. People use it to share their ideas, exchange information, and discuss thoughts. Given its prevalence and widespread reach, social media must remain a safe space for people. Content generated on social media can b...
- Whose Voice Counts? Mapping Stakeholder Perspectives on AI Through Public Submissions to the U.S. Government
As artificial intelligence (AI) systems become more common in our daily lives, it is important to understand how different stakeholders comprehend and envisage the role that these technologies play in shaping social, political, and economic realities. In this paper, we investigate public perceptions...
- Two is better than one: A Collapse-free Multi-Reward RLIF Training Framework
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of LLMs, but often depends on external supervision from human annotations or gold-standard solutions. Reinforcement learning from internal feedback (RLIF) has recently emerged as a scalable unsuper...