初始化项目,由ModelHub XC社区提供模型
Model: ICTNLP/UMA-4B Source: Original Platform
This commit is contained in:
102
README.md
Normal file
102
README.md
Normal file
@@ -0,0 +1,102 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
base_model:
|
||||
- Qwen/Qwen3-4B-Instruct-2507
|
||||
base_model_relation: finetune
|
||||
tags:
|
||||
- memory-agent
|
||||
- reinforcement-learning
|
||||
- long-context
|
||||
- tool-use
|
||||
- qwen3
|
||||
- grpo
|
||||
arxiv: 2602.18493
|
||||
---
|
||||
|
||||
# UMA-4B (Generalist)
|
||||
|
||||
**UMA-4B** is the Generalist checkpoint of the Unified Memory Agent (UMA) introduced in [Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning](https://arxiv.org/abs/2602.18493).
|
||||
|
||||
UMA is a tool-using memory agent that incrementally maintains a compact core summary and a structured key-value Memory Bank. The same policy performs memory construction and downstream question answering through explicit memory and retrieval operations.
|
||||
|
||||
- **Code:** [github.com/ictnlp/unified-memory-agent](https://github.com/ictnlp/unified-memory-agent)
|
||||
- **Paper:** [arXiv:2602.18493](https://arxiv.org/abs/2602.18493)
|
||||
- **Specialist checkpoint:** [ICTNLP/UMA-LedgerQA-4B](https://huggingface.co/ICTNLP/UMA-LedgerQA-4B)
|
||||
|
||||
## Checkpoint Variant
|
||||
|
||||
This repository contains the **Generalist** UMA checkpoint used for the paper's Test-Time Learning and Accurate Retrieval evaluations.
|
||||
|
||||
| Property | Value |
|
||||
| --- | --- |
|
||||
| Base model | `Qwen/Qwen3-4B-Instruct-2507` |
|
||||
| Parameters | 4B |
|
||||
| Weight format | BF16 Safetensors |
|
||||
| Training method | End-to-end reinforcement learning with Task-Stratified GRPO |
|
||||
| Training data | HotpotQA and the Mem-alpha corpus |
|
||||
| Ledger-QA training data | None |
|
||||
| Reported default context budget | 16K |
|
||||
|
||||
The Generalist and Specialist checkpoints share the same UMA architecture and tool interface. The Specialist checkpoint is additionally adapted to Ledger-QA; use this Generalist checkpoint for the broader cross-task setting.
|
||||
|
||||
## Intended Use
|
||||
|
||||
This checkpoint is intended for research on:
|
||||
|
||||
- long-context and streaming memory agents;
|
||||
- proactive structured memory construction;
|
||||
- memory maintenance with explicit tool calls;
|
||||
- downstream question answering over reusable memory;
|
||||
- evaluation and extension of the UMA framework.
|
||||
|
||||
The checkpoint is designed to run inside the UMA two-phase agent loop. A plain text-generation call loads the language model, but does not by itself instantiate the Memory Bank, retrieval tools, prompts, or memory-to-QA workflow.
|
||||
|
||||
## Loading the Weights
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_id = "ICTNLP/UMA-4B"
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
model_id,
|
||||
torch_dtype=torch.bfloat16,
|
||||
device_map="auto",
|
||||
)
|
||||
```
|
||||
|
||||
The model can also be served through an OpenAI-compatible inference server:
|
||||
|
||||
```bash
|
||||
vllm serve ICTNLP/UMA-4B \
|
||||
--max-model-len 16384 \
|
||||
--gpu-memory-utilization 0.8
|
||||
```
|
||||
|
||||
For full memory-agent inference, including the Memory Bank, memory tools, embedding retrieval, prompts, and benchmark runners, follow the [official repository](https://github.com/ictnlp/unified-memory-agent).
|
||||
|
||||
## Limitations
|
||||
|
||||
- The checkpoint is a research model and may generate incorrect answers or perform incorrect memory updates.
|
||||
- Agent behavior depends on the UMA prompt templates, tool implementations, retrieval backend, chunking policy, and inference configuration.
|
||||
- The model was primarily trained and evaluated on English-language research benchmarks.
|
||||
- Persistent-memory applications can involve sensitive information. Deployments should provide appropriate privacy controls, retention policies, and user oversight.
|
||||
- This checkpoint should not be used as the sole basis for high-stakes decisions.
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@article{zhang2026learning,
|
||||
title = {Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning},
|
||||
author = {Zhang, Kehao and Gui, Shangtong and Yang, Sheng and Chen, Wei and Feng, Yang},
|
||||
journal = {arXiv preprint arXiv:2602.18493},
|
||||
year = {2026}
|
||||
}
|
||||
```
|
||||
|
||||
Reference in New Issue
Block a user