fromtransformersimportAutoModelForCausalLM,AutoTokenizerimporttorch# Load the modelmodel=AutoModelForCausalLM.from_pretrained("EganAI/Qwen3-4B-Thinking-2507-20250813-033307-1",torch_dtype=torch.float16,device_map="auto")tokenizer=AutoTokenizer.from_pretrained("EganAI/Qwen3-4B-Thinking-2507-20250813-033307-1")# Example: MMLU-style questionprompt='''Question: The study of the distribution and determinants of health and disease in populations is:
A) Epidemiology
B) Ecology
C) Etiology
D) Endocrinology
Answer:'''inputs=tokenizer(prompt,return_tensors="pt")outputs=model.generate(**inputs,max_length=150,temperature=0.7,do_sample=True)response=tokenizer.decode(outputs[0],skip_special_tokens=True)print(response)
Inference with vLLM
fromvllmimportLLM,SamplingParamsllm=LLM(model="EganAI/Qwen3-4B-Thinking-2507-20250813-033307-1")sampling_params=SamplingParams(temperature=0.7,top_p=0.95,max_tokens=256)prompts=["Question: Explain quantum entanglement in simple terms."]outputs=llm.generate(prompts,sampling_params)
Technical Details
This model uses a layer-wise merging approach where each transformer layer is selected from different source models based on optimization criteria. This technique allows combining strengths from multiple fine-tuned models.
Merging Process
Layer Selection: Each layer (0-35 for this architecture) is independently selected from one of the source models
Non-layer Weights: Embeddings and final layers are taken from the base model
Optimization: The configuration was found through systematic optimization on the target benchmark
Limitations
This is an experimental merge and performance may vary on tasks outside the optimization targets
The model inherits limitations from its source models
Performance on general tasks may differ from benchmark scores
Citation
If you use this model, please cite the original source models and egan.ai
Note
This model is provided for research purposes. Always validate performance on your specific use case before deployment.