VOL. I NO. 2
PRICE: $0.00 / FREE

The Latent Chronicle

FRIDAY, JULY 31, 2026
FORECAST: HIGH COMPRESSION RATE, DENSE CONVOLUTIONS
Daily Dispatches & Bulletins
Nº 8
cs.CV▲ 87 UPVOTES

DISTILLALIGN REVOLUTIONIZES VIDEO GENERATION THROUGH DISTRIBUTIONAL HARMONIZATION

Paper: DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
By: Jiaxing Li, Kai Zou et al.

Researchers demonstrate that autoregressive video distillation suffers when initialization and matching stages pursue conflicting distributions. By introducing a latent space evaluation protocol and a joint distillation method combining reverse KL with consistency constraints, the technique successfully prevents coverage collapse and significantly boosts generation diversity.

Nº 11
cs.AI▲ 75 UPVOTES

METIS FOUNDATION MODEL INTRODUCES NATIVE PERSISTENT MEMORY FOR AUTONOMOUS AGENTS

Paper: Metis: Memory Foundation Model
By: Zeyu Zhang, Ziliang Guo et al.

Breaking away from reliance on external retrieval modules, this novel architecture embeds a dynamically evolving memory state directly into the neural backbone. Through gradient free online maintenance and memory attention compression, the system retains historical context across computation steps entirely within a frozen weight framework.

Nº 13
cs.LG▲ 76 UPVOTES

FRONTIS MA1 DRIVES RECURSIVE SELF IMPROVEMENT ACROSS MACHINE LEARNING ENGINEERING

Paper: Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
By: Junlin Yang, Che Jiang et al.

The OpenMLE framework combines verifiable task environments, operator learning, and long horizon search to study autonomous artificial intelligence improvement loops.

By aligning models around draft, improve, debug, and crossover operators, the thirty five billion parameter meta evolution agent substantially outperforms baseline architectures on complex benchmarks.

Nº 14
cs.CL▲ 78 UPVOTES

COUNTERFACTUAL REPLAY REDISTRIBUTES TOKEN LEVEL CREDIT IN RUBRIC GUIDED REINFORCEMENT LEARNING

Paper: CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
By: Bo-Wen Zhang, Junwei He et al.

Traditional policy optimization relies on response level rewards that fail to allocate credit to specific token spans. This new approach computes likelihood contrasts between rubric conditioned and criteria free prompts to generate normalized weights, effectively distributing graded reinforcement signals without requiring auxiliary scoring models.

Nº 20
cs.AI▲ 59 UPVOTES

ASKCHEM ESTABLISHES CLAIM CENTERED INFRASTRUCTURE FOR CHEMICAL LITERATURE SYNTHESIS

Paper: AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
By: Bing Yan, Gregory Wolfe et al.

This platform transforms scientific search by converting research papers into atomic, provenance carrying claims linked through an evidence graph and faceted taxonomy.

Grounded readers utilizing this architecture achieve flawless DOI resolution and superior citation density when assembling cross paper chemical knowledge.

Nº 24
cs.CV▲ 55 UPVOTES

PhiZero: A World Model Built Around Physical Language

By: Shuyao Shang, Yuqi Wang et al.

We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors.

Nº 34
cs.CL▲ 35 UPVOTES

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

By: Qixun Wang, Yang Shi et al.

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE).

Nº 36
cs.CL▲ 34 UPVOTES

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

By: Pengyu Wang, Benfeng Xu et al.

Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that varies corpus size along 28 strictly nested tiers spanning roughly 450-fold, while holding questions and a fixed bedrock of relevant and adversarial documents unchanged.

Nº 37
cs.CL▲ 32 UPVOTES

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

By: Yu Wang, Yi-Kai Zhang et al.

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate.

Nº 38
cs.CV▲ 29 UPVOTES

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

By: Haodong Li, Tianfei Ren et al.

Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process.

Nº 40
cs.CL▲ 28 UPVOTES

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

By: Yuhang Zhu, Mingxuan Du et al.

Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement.

Nº 46
cs.CV▲ 25 UPVOTES

Flux-OPD: On-Policy Distillation with Evolving Contexts

By: Yuran Wang, Zekun Wang et al.

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance.

Nº 48
cs.CL▲ 22 UPVOTES

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

By: Zhiyuan Yao, Yuxin Chen et al.

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution.

Nº 49
cs.CV▲ 22 UPVOTES

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

By: Hanzhang Zhou, Panrong Tong et al.

GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort.

Nº 52
cs.CV▲ 19 UPVOTES

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

By: Tengfei Liu, Yang Shi et al.

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task.

Nº 53
cs.CV▲ 18 UPVOTES

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

By: Yang Zhou, Zixuan Huang et al.

Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions.

Nº 54
cs.CL▲ 20 UPVOTES

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

By: Yihao Chen, Shi Chang et al.

Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks.

Nº 56
cs.CL▲ 16 UPVOTES

SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch

By: Yihao Chen, Shi Chang et al.

LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances.

Nº 57
cs.CL▲ 2 UPVOTES

Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems

By: Ali Zahid Raja

Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execution traces, tool calls, evidence links - and source-reliability estimation is long established (truth discovery, reputation systems).

Nº 58
cs.CL▲ 9 UPVOTES

Can Large Language Models Execute Parent Orders?

By: Zane Shen, Xinli Xu et al.

Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not hold in practice, or require task-specific training that limits adaptability to new settings.

Nº 59
cs.CL▲ 10 UPVOTES

MemHarness: Memory Is Reconstructed, Not Replayed

By: Rong Wu, Daocheng Fu et al.

Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent's current situation.

Nº 61
cs.CV▲ 2 UPVOTES

πR^2: Reactive Real-time Flow Policies

By: Sungjae Park, Shubham Tulsiani

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing reactivity.