MEMORY, ON DEMAND.

MemPilot

Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

The right evidence. The right model. The right amount of compute. A learned policy that puts agent memory on demand.

arXiv:2610.06830/Cite this work

A LEARNED POLICY. AN ADAPTIVE LOOP.
The MemPilot orchestrator reads the query and accumulated evidence. It chooses RETRIEVE or CURATE, sets the retrieval query and evidence count, and controls the curation instruction, selected LLM/VLM, and visual access. Both tools return text evidence to the next policy decision; the policy can also finish and answer. Expand
TRAINED TO BALANCEQuality ↑Cost ↓Latency ↓

Haozhen Zhang1Haodong Yue2Quanyu Long1Jianzhu Bao1Qingyuan Liu1Tao Feng3Bohan Liu1Weida Liang1Wenya Wang1

1 Nanyang Technological
University
2 Tsinghua
University
3 University of Illinois
Urbana-Champaign

Memory should
follow the question.

Before a question arrives, it is hard to know which details will matter. Compressing everything in advance can be expensive—and can lose exactly the evidence an agent later needs.

MemPilot keeps both a compact memory bank and the original multimodal history. At query time, a learned policy decides what to retrieve, how to curate it, and when enough evidence has been collected.

01 / ACCESS

Two views of memory.

Retrieve efficiently from query-agnostic memory. Revisit original text and images when the question needs more.

02 / DELEGATE

The model fits the task.

Choose from heterogeneous LLMs and VLMs, with control over evidence count, instructions, and visual access.

03 / ADAPT

Every step has a purpose.

Learn to balance answer quality, cost, and latency, with credit assigned to the evidence each step contributes.

Read the abstract

Memory has become integral to the LLM agent ecosystem, supporting information retention and reuse across interactions. However, most existing agent memory systems construct memory in a query-agnostic manner, which can incur unnecessary preprocessing cost and discard details that later prove essential. Recent studies have begun shifting memory processing toward runtime adaptation, but typically specialize in particular operations or fixed processing schemes, leaving flexible control over performance, cost, and latency largely underexplored.

To address this challenge, we present MemPilot, a flexible framework that orchestrates on-demand memory curation under different performance–cost–latency preferences. Specifically, we optimize a multi-step LLM policy via reinforcement learning to iteratively choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. The policy jointly controls evidence amount, curation instructions, model selection, and visual access, enabling fine-grained allocation of runtime computation. To optimize this policy under competing objectives, we adapt objective-wise advantage decoupling by separately estimating each objective’s advantage before aggregation. Moreover, we introduce prefix-based marginal utility estimation for fine-grained credit assignment across multi-step rollouts. Experiments on five multimodal agent-memory benchmarks demonstrate favorable performance–cost–latency trade-offs across optimization preferences, with preference sweeps yielding broader frontiers than existing trade-off-aware baselines.

A policy at the helm.

Retrieve. Curate. Reassess.
Continue until the evidence is enough.

A flexible decision loop: choose either action, repeat as needed, and run independent operations in parallel.

ACTION / RETRIEVE

Useful memory, without another model call.

Search the query-agnostic memory bank with a learned retrieval query and evidence count. The bank can come from an existing memory system; retrieved text returns directly to the orchestrator.

Retrieval query rEvidence count k

The default bank uses LLMLingua-2. Raw text and images remain available for later curation.

HOW THE POLICY LEARNS

Reward the outcome.
Credit the useful steps.

Separate the objectives

Normalize quality, cost, and latency advantages independently, then combine them using the chosen preference weights.

Measure marginal utility

The policy probes an answer before and after each memory-gathering stage. The change in answer quality provides a learning signal for that stage.

Paper architecture: the orchestrator selects RETRIEVE or CURATE; factorized controls govern evidence, instructions, model selection and visual access. Training combines objective-wise advantage decoupling with prefix-based marginal utility. Expand figure
FIG. 01 On-demand memory orchestration and resource-aware training.Paper, §3

Better answers.
Room to choose.

Five multimodal benchmarks. Two answer models.
Explore quality and cost at three operating points.

Mem-Gallery

In-domain evaluation
LLM-JUDGE (%) ↑COST (USD/SAMPLE) ↓
A-Mem Best baseline
51.45
$0.028
MemPilot Perf
65.82
$0.043
MemPilot Bal
64.27
$0.014
MemPilot Cost
39.09
$0.0082

Baseline selected by highest LLM-Judge score. All methods below.

All methods & metrics 11 methods
Mem-Gallery · Qwen3-VL-4B-Instruct · Paper, Table 1
MethodF1 ↑LLM-Judge ↑Cost (USD/sample) ↓
A-Mem54.2451.45$0.028
SimpleMem27.3232.00$0.019
MIRIX36.7049.45$0.33
M2A39.0550.73$0.83
MemVerse30.9244.27$0.47
GAM43.4646.36$0.015
AgeMem38.3839.36$0.03
Mem-T36.1834.55$0.019
MemPilot-Perf53.7365.82$0.043
MemPilot-Bal57.2664.27$0.014
MemPilot-Cost32.6739.09$0.0082

Bold cells indicate the best result in each column. F1 and LLM-Judge are percentages; higher is better. Lower cost is better.

Source: Table 1For fair comparison with baselines, final answers use the same answer model across methods. Cost follows the paper’s USD/sample accounting, normalized by unique dataset–conversation pairs (Appendix A.4). MemEye and MEMLENS are evaluated without training on either benchmark. Relative gains and cost savings use rounded table values.

One framework.
A broader frontier.

Vary the optimization preference to choose a policy
that fits the quality and resources you need.

Original Figure 2. MemPilot traces a broader performance–cost and performance–latency frontier than the compared baselines. The cost axis contains a break; latency is a proxy in seconds per query. Expand figure
FIG. 02 Macro-averaged over five benchmarks. Lines connect Pareto-optimal points.Paper, §4.4

More room to allocate compute.

Preference sweeps cover lightweight through performance-oriented policies, extending the trade-off space beyond fixed operating points.

Cost and latency both matter.

A cheaper policy is not necessarily a faster one. MemPilot optimizes both.

Different priorities.
Different decisions.

The policy changes how it gathers evidence,
how often it looks at images, and which models it calls.

RUNTIME COMPUTATION

See how the policy adapts.

Quality-orientedLower cost
Cost (USD/sample) · five-benchmark averages
Cost (USD/sample)$0.0409$0.0380$0.0108$0.0059
Policy steps2.062.011.901.00
CURATE actions54.3%51.5%49.8%0.0%
Evidence count3.842.884.003.00
Image access24.7%27.5%0.0%0.0%
Policy output tokens184.8198.2142.948.0

Source: Figure 3Four separately optimized policies, ordered by resource usage. These are reported operating points, not a live inference-time slider.

ABLATION STUDY

The useful steps matter.

Removing delegation or fine-grained credit assignment substantially reduces answer quality.

MemPilot-Bal
37.40
Without GDPO
21.94
Without retrieval–instruction decoupling
15.35
Without marginal utility
14.82
Without LLM/VLM delegation
14.58

LLM-Judge (%) · axis 0–40 · five-benchmark average
Qwen3-VL-4B · Table 2

Ablated variants also cost less ($0.0054–$0.0059 vs. $0.011); this is not a matched-cost comparison.

MEMORY-BANK COMPATIBILITYNo retraining

Different banks.
One learned policy.

Swap in LLMLingua-2, SimpleMem, or A-Mem memories. The same orchestrator continues to control how evidence is gathered.

Figure 4 from the paper. MemPilot with its default LLMLingua-2 bank occupies the lower-cost regime, SimpleMem an intermediate regime, and A-Mem a higher-cost regime with a higher performance ceiling. The orchestrator is not retrained. Expand figure
FIG. 04 Five-benchmark average · Qwen3-VL-4B

Compatible, with different trade-offs. The default bank offers a lower-cost regime; A-Mem reaches a higher performance ceiling at higher cost.

Use your own memory bank

Take memory
further.

Explore the implementation, reproduce the results,
or bring your own memory bank.

Explore the code
Reference / BibTeX
@article{zhang2026mempilot,
  title={MemPilot: Orchestrating On-Demand
         Multimodal Memory Curation for LLM Agents},
  author={Zhang, Haozhen and Yue, Haodong and
          Long, Quanyu and Bao, Jianzhu and
          Liu, Qingyuan and Feng, Tao and
          Liu, Bohan and Liang, Weida and Wang, Wenya},
  journal={arXiv preprint arXiv:2610.06830},
  year={2026}
}

arXiv:2610.06830 [cs.CL] · 2026

Paper figure

Open original ↗