Two views of memory.
Retrieve efficiently from query-agnostic memory. Revisit original text and images when the question needs more.
MEMORY, ON DEMAND.
Orchestrating On-Demand Multimodal Memory Curation for LLM Agents
The right evidence. The right model. The right amount of compute. A learned policy that puts agent memory on demand.
Before a question arrives, it is hard to know which details will matter. Compressing everything in advance can be expensive—and can lose exactly the evidence an agent later needs.
MemPilot keeps both a compact memory bank and the original multimodal history. At query time, a learned policy decides what to retrieve, how to curate it, and when enough evidence has been collected.
Retrieve efficiently from query-agnostic memory. Revisit original text and images when the question needs more.
Choose from heterogeneous LLMs and VLMs, with control over evidence count, instructions, and visual access.
Learn to balance answer quality, cost, and latency, with credit assigned to the evidence each step contributes.
Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored.
To address this challenge, we present MemPilot, a flexible framework that orchestrates on-demand memory curation under different performance–cost–latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective’s advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance–cost–latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.
Retrieve. Curate. Reassess.
Continue until the evidence is enough.
A flexible decision loop: choose either action, repeat as needed, and run independent operations in parallel.
Search the query-agnostic memory bank with a learned retrieval query and evidence count. The bank can come from an existing memory system; retrieved text returns directly to the orchestrator.
The default bank uses LLMLingua-2. Raw text and images remain available for later curation.
Normalize quality, cost, and latency advantages independently, then combine them using the chosen preference weights.
The policy probes an answer before and after each memory-gathering stage. The change in answer quality provides a learning signal for that stage.

Five multimodal benchmarks. Two answer models.
Explore quality and cost at three operating points.
| Method | F1 ↑ | LLM-Judge ↑ | Cost (USD/sample) ↓ |
|---|---|---|---|
| A-Mem | 54.24 | 51.45 | $0.028 |
| SimpleMem | 27.32 | 32.00 | $0.019 |
| MIRIX | 36.70 | 49.45 | $0.33 |
| M2A | 39.05 | 50.73 | $0.83 |
| MemVerse | 30.92 | 44.27 | $0.47 |
| GAM | 43.46 | 46.36 | $0.015 |
| AgeMem | 38.38 | 39.36 | $0.03 |
| Mem-T | 36.18 | 34.55 | $0.019 |
| MemPilot-Perf | 53.73 | 65.82 | $0.043 |
| MemPilot-Bal | 57.26 | 64.27 | $0.014 |
| MemPilot-Cost | 32.67 | 39.09 | $0.0082 |
Bold cells indicate the best result in each column. F1 and LLM-Judge are percentages; higher is better. Lower cost is better.
Source: Table 1For fair comparison with baselines, final answers use the same answer model across methods. Cost follows the paper’s USD/sample accounting, normalized by unique dataset–conversation pairs (Appendix A.4). MemEye and MEMLENS are evaluated without training on either benchmark. Relative gains and cost savings use rounded table values.
Vary the optimization preference to choose a policy
that fits the quality and resources you need.

Preference sweeps cover lightweight through performance-oriented policies, extending the trade-off space beyond fixed operating points.
A cheaper policy is not necessarily a faster one. MemPilot optimizes both.
The policy changes how it gathers evidence,
how often it looks at images, and which models it calls.
| Cost (USD/sample) | $0.0409 | $0.0380 | $0.0108 | $0.0059 |
|---|---|---|---|---|
| Policy steps | 2.06 | 2.01 | 1.90 | 1.00 |
| CURATE actions | 54.3% | 51.5% | 49.8% | 0.0% |
| Evidence count | 3.84 | 2.88 | 4.00 | 3.00 |
| Image access | 24.7% | 27.5% | 0.0% | 0.0% |
| Policy output tokens | 184.8 | 198.2 | 142.9 | 48.0 |
Source: Figure 3Four separately optimized policies, ordered by resource usage. These are reported operating points, not a live inference-time slider.
Removing delegation or fine-grained credit assignment substantially reduces answer quality.
LLM-Judge (%) · axis 0–40 · five-benchmark average
Qwen3-VL-4B · Table 2
Ablated variants also cost less ($0.0054–$0.0059 vs. $0.011); this is not a matched-cost comparison.
Swap in LLMLingua-2, SimpleMem, or A-Mem memories. The same orchestrator continues to control how evidence is gathered.
Compatible, with different trade-offs. The default bank offers a lower-cost regime; A-Mem reaches a higher performance ceiling at higher cost.
Use your own memory bankExplore the implementation, reproduce the results,
or bring your own memory bank.
@article{zhang2026mempilot,
title={MemPilot: Orchestrating On-Demand
Multimodal Memory Curation for LLM Agents},
author={Zhang, Haozhen and Yue, Haodong and
Long, Quanyu and Bao, Jianzhu and
Liu, Qingyuan and Feng, Tao and
Liu, Bohan and Liang, Weida and Wang, Wenya},
journal={arXiv preprint arXiv:2610.06830},
year={2026}
}