Instructions to use anshumanatrey/pharmarl-llama-3b-trained-anshuman with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use anshumanatrey/pharmarl-llama-3b-trained-anshuman with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Llama-3.2-3B-Instruct") model = PeftModel.from_pretrained(base_model, "anshumanatrey/pharmarl-llama-3b-trained-anshuman") - Notebooks
- Google Colab
- Kaggle
PharmaRL β Llama-3.2-3B-Instruct trained via GRPO
LoRA adapter trained on top of meta-llama/Llama-3.2-3B-Instruct using GRPO (Group Relative Policy Optimization) inside the PharmaRL OpenEnv-native chemistry environment.
The model learns to design drug-like molecules step by step by emitting JSON molecular edits (add fragment, remove, substitute atom, terminate) against composite chemistry rewards: Lipinski compliance, QED drug-likeness, synthetic accessibility, and TDC bioactivity classifiers (DRD2, GSK3B, JNK3).
Headline numbers
- Parse rate: 100% from step 0 (chat-template prompting)
- Final mean reward: +2.079 (easy-tier curriculum, step 199)
- Peak max reward: +8.842 (step 96)
- Best molecule emitted (step 150):
CC1C(C(=O)O)C(c2ccccc2)CCN1Cc1ccncc1β QED 0.94, non-trivial docking signal - Training: 200 GRPO steps, group size G=8, single NVIDIA A10G GPU, ~6 hours wall clock
Training stack
- Base: Llama-3.2-3B-Instruct (Meta, September 2024)
- PEFT: LoRA, rank 16, alpha 32, all linear projections
- Algorithm: GRPO β group-standardized advantage with K3 KL anchor
- Hyperparameters: lr=5e-6, KL Ξ²=0.04, clip Ξ΅=0.2, max episode steps=20
How to use
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "meta-llama/Llama-3.2-3B-Instruct"
adapter = "anshumanatrey/pharmarl-llama-3b-trained-anshuman"
tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, adapter)
Run inside the PharmaRL environment for full episode rollouts against the chemistry oracles.
Links
- Environment (HF Space): anshumanatrey/pharmarl
- Code: github.com/AnshumanAtrey/pharmarl
- Run audit bundle: runs/anshuman-a10g-llama3b
- Research paper (Zenodo): zenodo.org/records/19788570
Submission context
Submitted to the Meta PyTorch OpenEnv Hackathon β Round 2 Grand Finale, Bangalore, April 25β26, 2026.
Team AI Mafias β Anshuman Atrey, Sahil Shah, Vijay Kota.
Limitations
- 3B base, hackathon-scale training (200 GRPO steps, single GPU, ~6 hours wall clock).
- Composite reward is reward-hackable in the sense Renz et al. (2020) document β see the molecule audit trail in the run bundle for an example of mid-training reward exploitation at step 100 and recovery at step 150.
- No JNK3 held-out evaluation completed within the hackathon time budget; the env supports it but the eval pass is pending.
- This is infrastructure for LLM-as-policy chemistry research, not a clinical or production drug-discovery tool. The model does not produce candidate drugs.
Citation
@misc{atrey2026pharmarl,
title = {AI Alchemy in Medicine: A Vision for LLM-as-Policy Molecular Design via {OpenEnv}},
author = {Atrey, Anshuman and Shah, Sahil and Kota, Vijay},
year = {2026},
publisher = {Zenodo},
url = {https://zenodo.org/records/19788570},
note = {Meta PyTorch OpenEnv Hackathon, Round 2 Grand Finale, Bangalore}
}
- Downloads last month
- 10
Model tree for anshumanatrey/pharmarl-llama-3b-trained-anshuman
Base model
meta-llama/Llama-3.2-3B-Instruct