AI News Archive: July 27, 2026 — Part 22
Sourced from 500+ daily AI sources, scored by relevance.
- A Modern ConvNet for Solar Filament Detection
Automated solar filament detection using deep learning faces several challenges. Semantic segmentation of solar filaments is a complicated multiscale feature extraction task with long-tail distribution. Furthermore, a large-scale, highly complete, and finely detailed dataset has become mandatory for...
- FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and referenc...
- Rethinking Expert Training for Model Merging with Prompt Learning
Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approaches largely focus on improving the merging procedure itself and typically assume experts obtained through full-parameter fine-tuning. In this work,...
- RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models
On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. However, existing methods often rely on verified solution traces, explanations generated by external models, or manually loc...
- MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning
Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features can discard text, local attributes, and spatial relationships, while high-resolution inputs substantially increase context length and inference latenc...
- EasyVoice
Clone any voice in 3 seconds — locally, in 10 languages
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing, text-based methods rely on coordinates or index prediction, severely limiting the perceptual capabilities of the model...
- Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG
Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data from remote deployments. This paper proposes a low-consumption underwater monitoring architecture that combines always-on edge sensing with selective hi...
- Space CRM
Unified AI platform for Email, LinkedIn & WhatsApp.
- VoiceID
Your Voice. Your Identity.
- Tiune
Dictation at the speed of thought
- eva
为shell而生的单文件智能体,无依赖,随处运行
- Stress-Test Your Product
Practical stress-test cards for AI and no-code founders
- Docsphere Solutions
AI-powered academic research, simplified
- Neurix Reader
Read a book, watch your notes write themselves
- Lumi Language
AI language tutor for personal goals
- Terriqon
Collect field data offline. Get AI reports in minutes.
- ResumeAI
Free AI Resume Optimizer — Upload, paste, get matched
- PromptFlow — AI Prompt Pack + Automation
100+ AI prompts + 22 automation workflows.
- Plain AI Office
Turn Markdown into live decks, docs, and dashboards
- Scholarship Copilot
Find direct targeted scholarship matches in minutes
- Ai pin maker
让设计变简单
- CodemandiA
Detect and explain code errors effortlessly using AI
- Codaiman
Building the future with AI-powered solutions.
- Legalysis
Understand any contract before you sign it
- VoiceDrop
AI Ringless Voicemail for Smarter Sales Outreach
- Al Business Operations Blueprint
Automate your content & business workflows with Al in Notion
- 实在agent
自主操作软件的企业级智能办公助理。
- AIVM Launch
The governed AI brain for people and agents.
- PixelPromptX
Create Stunning AI Images. Instantly.
- Multi-Agent AI for BFSI Support
Resolution-first multi-agent AI for BFSI customer support
- MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving
Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a ...
- MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale
We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of consumer-grade microscope close-ups, MicroZoom synthesizes a seamless, gigapixel-resolution image grounded in the material character of the real refe...
- DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement
With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches still struggle to jointly preserve style fidelity, geometric consistency, and generation efficiency, as most of them sti...
- SADe: Sparse-Atom Support Decontamination for Few-Shot Segmentation with Weak Support Annotations
Few-shot segmentation (FSS) commonly assumes clean pixel-level support masks, yet practical support supervision often uses boxes, scribbles, coarse masks, or pseudo-masks. These weak annotations may include texture-similar distractors and background context alongside the target, contaminating class ...
- Panda: Unsupervised Pelvic Anomaly Detection for Real-Time MR Imaging
Female pelvic diseases remain an under researched area characterized by often delayed diagnosis. While pelvic MRI offers superior soft-tissue contrast for diagnosis and image-guided procedures, real-time anomaly detection remains challenging due to physiological motion, tissue deformation, and instr...
- MMOE: Modernizing Diffusion Transformers with Efficient Expert Design
Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly...
- QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment
Video instance segmentation (VIS) requires models to detect, segment, and track object identities across frames, and most methods enforce temporal consistency through video-level supervision. Image-only training approaches, with MinVIS as one prominent example, have challenged this assumption, reach...
- CameraAnything: Refilming Videos with Arbitrary Camera Control
We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrin...
- Denoising 3D images: robustness of persistent homology measures
When computing sub/super-level-set persistent homology (PH), the effect of noise may introduce millions of (short-lived) topological generators, presenting an obstacle to both the computation of PH of large 3D images, and any analysis of PH that incorporates the number of generators. As such, it is ...
- NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction
Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their depth quality. SLAM systems can benefit greatly from such strong depth sensing, where reliable geometry enables stable tracking and faithful reconstructio...
- MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE becomes a primary bottleneck. Existing approaches largely follow two paradigms: sparse voxel-based representations achieve strong reconstruction qualit...
- InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting
Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety. Even though current methods have achieved good performance, the strong spatial-temporal modeling between the input mul...
- IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data
This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2026 International Joint Conference on Biometrics (IJCB 2026). The competition received a total of eight valid submissions from four distinct teams a...
- Accuracy potential of visual localization exploiting high-end street-level imagery
Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, surveying, robotics, and augmented and mixed reality. Visual localization can serve as a complementary positioning modality to GNSS, whose applicabilit...
- GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion
Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which be...
- Agentica
Automated AI-agent investing across global tokenized assets
- HistoGPA: A Context-Conditioned Gene-Prior Attention Framework for Histology-Based Spatial Gene Expression Prediction
Predicting spatial gene expression from routine hematoxylin and eosin (H&E) images provides a practical complement to experimental spatial transcriptomics. Existing approaches focus on local or multi-scale visual features and often treat pretrained gene representations as fixed priors, although the ...
- TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across successive segments. A bounded cache retains local motion and phonetic context but discards older evidence, whereas attending...
- PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation
Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to prompt formulation. Existing prompt optimization methods mainly rely on text-side rewriting, prompt expansion, or external reward signals, offering...