19 lines
811 B
Markdown
19 lines
811 B
Markdown
---
|
|
library_name: transformers
|
|
license: apache-2.0
|
|
---
|
|
|
|
# E2L ALFWorld cold-start checkpoint (Qwen/Qwen3-8B)
|
|
|
|
Stage-1 cold-start checkpoint from
|
|
[Ch1nyzzz/e2-learning](https://github.com/Ch1nyzzz/e2-learning): full-parameter
|
|
online mistake-driven experience learning (next-observation SFT gated by a
|
|
semantic judge) on ALFWorld `AlfredTWEnv`, initialized from
|
|
`Qwen/Qwen3-8B`. Source run artifact: `outputs/alfworld_qwen3_8b_active_10000_hf/checkpoints/env_step_006220`.
|
|
|
|
Intended use: initialization ("cold start") for Stage-2 policy GRPO training
|
|
with [verl-agent](https://github.com/langfengQ/verl-agent). Load as a regular
|
|
Hugging Face checkpoint (`AutoModelForCausalLM.from_pretrained`).
|
|
`controller_state.json` is experiment-controller provenance metadata and can
|
|
be ignored by inference/training stacks.
|