ICML 2026RESEARCH PROJECT

BudgetMem

Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory

The right memory. The right amount of compute.
A learned router that makes agent memory adapt to the query—and to the budget.

Discover the framework
MEMORY, IN THE MAKINGRUNTIME
Query + retrieved history
Shared budget-tier routerRL

Illustrative allocations. Actual tiers depend on the query.

Balanced example: Filter Mid, Entity High, Temporal Mid, Topic Low, Summary High.

Haozhen Zhang*,1Haodong Yue*,2Tao Feng3Quanyu Long1Jianzhu Bao1Bowen Jin3Weizhi Zhang4Xiao Li5Jiaxuan You3Chengwei Qin6Wenya Wang1

* Equal contribution

1Nanyang Technological
University
2Tsinghua
University
3University of Illinois
Urbana-Champaign
4University of Illinois
Chicago
5Sun Yat-sen
University
6HKUST
(Guangzhou)

Memory, on demand.
Compute, under control.

Agent memory should work for the question being asked. BudgetMem extracts memory at runtime, using the current query to decide what matters and how much computation to spend.

A modular pipeline exposes Low, Mid, and High budget tiers. A lightweight router learns to choose a tier for each module, balancing answer quality with extraction cost through reinforcement learning.

Read the abstract
01 /

Query-aware extraction

Retrieve relevant raw history and build memory when a query arrives. Keep computation focused on the evidence that matters.

02 /

Module-level control

Give every stage its own compute budget, with three tiers and a consistent interface across the memory pipeline.

03 /

Learned budget routing

Use the query and intermediate states to allocate compute, guided by a reward that accounts for both quality and cost.

A small router.
A more flexible memory.

5memory modules
3budget tiers per module
1shared learned router

Every module.
The right level of effort.

From raw history to a useful memory: retrieve, filter, extract entity, temporal, and topic context, then summarize. The router decides how each stage runs.

FIG. 01BudgetMem architectureExpand figure
BudgetMem retrieves raw chunks, filters them, extracts entity, temporal, and topic context in parallel, and summarizes the results. A shared reinforcement-learned router chooses Low, Mid, or High tiers for each module.
One modular pipeline. A shared router. Query-aware decisions at every stage.

Three ways to define a budget.

The same tier interface. Three complementary axes of control.

BUDGETMEM–IMP

Implementation

Change the method behind each module.

  1. LowLightweight heuristics
  2. MidTask-specific models
  3. HighLLM-based processing

Broad budget coverage, with rapid quality gains at moderate cost.

BUDGETMEM–REA

Reasoning

Change how much the model reasons.

  1. LowDirect inference
  2. MidChain-of-thought
  3. HighMulti-step / reflection

Fine-grained quality control within a more concentrated cost range.

BUDGETMEM–CAP

Capacity

Change the size of the model in use.

  1. LowSmall model
  2. MidMedium model
  3. HighLarge model

A wide compute range, extending quality as the budget grows.

Better memory.
Measurable gains.

Evaluated on three benchmarks with two model backbones. Explore the performance-first results, with answer quality and cost side by side.

LongMemEval

LLaMA-3.3-70B-Instruct · performance-first setting

LLM-Judge (%) ↑

BudgetMem variantsStrongest baseline by Judge

View all methods10 methods · F1, Judge & cost
LongMemEval · LLaMA-3.3-70B-Instruct · performance-first (λ = 0)
MethodF1 ↑Judge ↑Cost (USD) ↓
ReadAgent20.7527.7213.68
MemoryBank26.7432.673.94
A-MEM21.7433.1780.02
LangMem12.0017.0016.60
Mem027.7042.0813.57
MemoryOS12.9733.5038.83
LightMem26.7448.515.28
BudgetMem–IMP37.4756.000.71
BudgetMem–REA40.5358.000.67
BudgetMem–CAP40.2460.500.80

Source: Table 1 in the paper (arXiv v3). F1 and Judge are reported as percentages. Cost is the paper’s reported token-based cost for the evaluated split. Qwen results evaluate transfer of the router without retraining.

More control.
Across the cost curve.

Varying the cost weight λ traces performance–cost frontiers on LoCoMo. BudgetMem provides explicit control over the compute spent on memory extraction.

Implementation and capacity tiering cover a broader range of budgets. Reasoning tiering offers finer quality adjustments within a narrower cost band.

FIG. 02Performance–cost frontiersExpand figure
Three plots compare LLM-Judge scores against cost on LoCoMo. BudgetMem traces controllable frontiers across implementation, reasoning, and capacity tiering relative to memory-system baselines.
Different tiering strategies offer complementary ways to navigate the quality–cost trade-off.

What makes the routing work?

ABLATION STUDY

Balance the learning signals.

Task reward and cost reward can differ substantially in scale. Without reward-scale alignment, the router becomes overly conservative and favors the Low tier.

On LoCoMo with capacity tiering, alignment supports a more graded use of tiers and a smoother performance–cost frontier.

Balanced rewards → meaningful budget control
Ablation of reward-scale alignment: removing alignment causes conservative low-tier routing and lower Judge scores.
Reward-scale alignment · LoCoMo · Capacity tiering
ROUTING BEHAVIOR

A budget response you can see.

As the cost weight λ increases, the router shifts its selections from higher-cost tiers to cheaper ones.

The module-level selection ratios on LongMemEval make this behavior interpretable: compute allocation changes with cost pressure across the pipeline.

Higher cost pressure → more economical tiers
Low, Mid, and High tier selection ratios across modules on LongMemEval, shifting toward lower-cost tiers as the cost weight increases.
Budget-tier selection ratios · LongMemEval · Capacity tiering
RETRIEVAL SENSITIVITY

Enough evidence. Less noise.

More retrieved chunks increase cost, but quality does not improve monotonically. Too few chunks omit useful evidence; too many introduce redundant or weakly relevant context.

In the reported LoCoMo setting, retrieving five chunks offers the best balance between cost and quality.

Relevant context matters more than volume
Judge scores and cost at different retrieval sizes on LoCoMo; five retrieved chunks provide the best reported balance.
Retrieval-size sensitivity · LoCoMo · All three tiering strategies

Citation

If BudgetMem helps your research, please consider citing our paper.

citation.bib
@article{BudgetMem,
  title   = {Learning Query-Aware Budget-Tier Routing
             for Runtime Agent Memory},
  author  = {Haozhen Zhang and Haodong Yue and
             Tao Feng and Quanyu Long and Jianzhu Bao
             and Bowen Jin and Weizhi Zhang and
             Xiao Li and Jiaxuan You and Chengwei Qin
             and Wenya Wang},
  journal = {arXiv preprint arXiv:2602.06025},
  year    = {2026}
}

Original