--- license: apache-2.0 tags: - pytorch - gpt2 - molecule-compound pipeline_tag: text-generation --- # molcrawl-compounds-guacamol-gpt2-small ## Model Description GPT-2 small (124M parameters) fine-tuned on [GuacaMol](https://github.com/BenevolentAI/guacamol) SMILES data, starting from the `molcrawl-compounds-gpt2-small` pre-trained model. The tokenizer is a character-level BPE tokenizer (vocab_size=612). Input SMILES strings should be passed **without** spaces. The `[SEP]` token (id=13) is used as the end-of-sequence marker. ## Datasets - **GuacaMol**: [https://github.com/BenevolentAI/guacamol](https://github.com/BenevolentAI/guacamol) (Fine-tuning dataset) - **Model Type**: gpt2 - **Data Type**: Molecule/Compound - **Training Date**: 2026-04-24 ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch model = AutoModelForCausalLM.from_pretrained("kojima-lab/molcrawl-compounds-guacamol-gpt2-small") tokenizer = AutoTokenizer.from_pretrained("kojima-lab/molcrawl-compounds-guacamol-gpt2-small") # Generate SMILES string prompt = "CC(=O)O" input_ids = tokenizer.encode(prompt, return_tensors="pt") with torch.no_grad(): output_ids = model.generate( input_ids, max_new_tokens=50, do_sample=True, temperature=0.8, eos_token_id=tokenizer.convert_tokens_to_ids("[SEP]"), # [SEP] is EOS for compounds pad_token_id=0, ) print(tokenizer.decode(output_ids[0], skip_special_tokens=True)) ``` ## Source Code Training pipeline, configuration files, and data preparation scripts are available in the MolCrawl GitHub repository: [https://github.com/mmai-framework-lab/MolCrawl](https://github.com/mmai-framework-lab/MolCrawl) ## License This model is released under the APACHE-2.0 license. ## Citation If you use this model, please cite: ```bibtex @misc{molcrawl_compounds_guacamol_gpt2_small, title={molcrawl-compounds-guacamol-gpt2-small}, author={{RIKEN}}, year={2026}, publisher={{Hugging Face}}, url={{https://huggingface.co/kojima-lab/molcrawl-compounds-guacamol-gpt2-small}} } ```