3.8 KiB
license, language, library_name, pipeline_tag, base_model, base_model_relation, tags, arxiv
| license | language | library_name | pipeline_tag | base_model | base_model_relation | tags | arxiv | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 |
|
transformers | text-generation |
|
finetune |
|
2602.18493 |
UMA-4B (Generalist)
UMA-4B is the Generalist checkpoint of the Unified Memory Agent (UMA) introduced in Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning.
UMA is a tool-using memory agent that incrementally maintains a compact core summary and a structured key-value Memory Bank. The same policy performs memory construction and downstream question answering through explicit memory and retrieval operations.
- Code: github.com/ictnlp/unified-memory-agent
- Paper: arXiv:2602.18493
- Specialist checkpoint: ICTNLP/UMA-LedgerQA-4B
Checkpoint Variant
This repository contains the Generalist UMA checkpoint used for the paper's Test-Time Learning and Accurate Retrieval evaluations.
| Property | Value |
|---|---|
| Base model | Qwen/Qwen3-4B-Instruct-2507 |
| Parameters | 4B |
| Weight format | BF16 Safetensors |
| Training method | End-to-end reinforcement learning with Task-Stratified GRPO |
| Training data | HotpotQA and the Mem-alpha corpus |
| Ledger-QA training data | None |
| Reported default context budget | 16K |
The Generalist and Specialist checkpoints share the same UMA architecture and tool interface. The Specialist checkpoint is additionally adapted to Ledger-QA; use this Generalist checkpoint for the broader cross-task setting.
Intended Use
This checkpoint is intended for research on:
- long-context and streaming memory agents;
- proactive structured memory construction;
- memory maintenance with explicit tool calls;
- downstream question answering over reusable memory;
- evaluation and extension of the UMA framework.
The checkpoint is designed to run inside the UMA two-phase agent loop. A plain text-generation call loads the language model, but does not by itself instantiate the Memory Bank, retrieval tools, prompts, or memory-to-QA workflow.
Loading the Weights
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ICTNLP/UMA-4B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
The model can also be served through an OpenAI-compatible inference server:
vllm serve ICTNLP/UMA-4B \
--max-model-len 16384 \
--gpu-memory-utilization 0.8
For full memory-agent inference, including the Memory Bank, memory tools, embedding retrieval, prompts, and benchmark runners, follow the official repository.
Limitations
- The checkpoint is a research model and may generate incorrect answers or perform incorrect memory updates.
- Agent behavior depends on the UMA prompt templates, tool implementations, retrieval backend, chunking policy, and inference configuration.
- The model was primarily trained and evaluated on English-language research benchmarks.
- Persistent-memory applications can involve sensitive information. Deployments should provide appropriate privacy controls, retention policies, and user oversight.
- This checkpoint should not be used as the sole basis for high-stakes decisions.
Citation
@article{zhang2026learning,
title = {Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning},
author = {Zhang, Kehao and Gui, Shangtong and Yang, Sheng and Chen, Wei and Feng, Yang},
journal = {arXiv preprint arXiv:2602.18493},
year = {2026}
}