--- library_name: transformers license: apache-2.0 --- # E2L ALFWorld cold-start checkpoint (Qwen/Qwen3-8B) Stage-1 cold-start checkpoint from [Ch1nyzzz/e2-learning](https://github.com/Ch1nyzzz/e2-learning): full-parameter online mistake-driven experience learning (next-observation SFT gated by a semantic judge) on ALFWorld `AlfredTWEnv`, initialized from `Qwen/Qwen3-8B`. Source run artifact: `outputs/alfworld_qwen3_8b_active_10000_hf/checkpoints/env_step_006220`. Intended use: initialization ("cold start") for Stage-2 policy GRPO training with [verl-agent](https://github.com/langfengQ/verl-agent). Load as a regular Hugging Face checkpoint (`AutoModelForCausalLM.from_pretrained`). `controller_state.json` is experiment-controller provenance metadata and can be ignored by inference/training stacks.