初始化项目,由ModelHub XC社区提供模型
Model: erv1n/e2l-alfworld-qwen25-7b-coldstart Source: Original Platform
This commit is contained in:
18
README.md
Normal file
18
README.md
Normal file
@@ -0,0 +1,18 @@
|
||||
---
|
||||
library_name: transformers
|
||||
license: apache-2.0
|
||||
---
|
||||
|
||||
# E2L ALFWorld cold-start checkpoint (Qwen/Qwen2.5-7B-Instruct)
|
||||
|
||||
Stage-1 cold-start checkpoint from
|
||||
[Ch1nyzzz/e2-learning](https://github.com/Ch1nyzzz/e2-learning): full-parameter
|
||||
online mistake-driven experience learning (next-observation SFT gated by a
|
||||
semantic judge) on ALFWorld `AlfredTWEnv`, initialized from
|
||||
`Qwen/Qwen2.5-7B-Instruct`. Source run artifact: `outputs/alfworld_qwen25_7b_active_10000_hf/checkpoints/env_step_001200`.
|
||||
|
||||
Intended use: initialization ("cold start") for Stage-2 policy GRPO training
|
||||
with [verl-agent](https://github.com/langfengQ/verl-agent). Load as a regular
|
||||
Hugging Face checkpoint (`AutoModelForCausalLM.from_pretrained`).
|
||||
`controller_state.json` is experiment-controller provenance metadata and can
|
||||
be ignored by inference/training stacks.
|
||||
Reference in New Issue
Block a user