PharmaRL β€” Llama-3.2-3B-Instruct trained via GRPO

LoRA adapter trained on top of meta-llama/Llama-3.2-3B-Instruct using GRPO (Group Relative Policy Optimization) inside the PharmaRL OpenEnv-native chemistry environment.

The model learns to design drug-like molecules step by step by emitting JSON molecular edits (add fragment, remove, substitute atom, terminate) against composite chemistry rewards: Lipinski compliance, QED drug-likeness, synthetic accessibility, and TDC bioactivity classifiers (DRD2, GSK3B, JNK3).

Headline numbers

  • Parse rate: 100% from step 0 (chat-template prompting)
  • Final mean reward: +2.079 (easy-tier curriculum, step 199)
  • Peak max reward: +8.842 (step 96)
  • Best molecule emitted (step 150): CC1C(C(=O)O)C(c2ccccc2)CCN1Cc1ccncc1 β€” QED 0.94, non-trivial docking signal
  • Training: 200 GRPO steps, group size G=8, single NVIDIA A10G GPU, ~6 hours wall clock

Training stack

  • Base: Llama-3.2-3B-Instruct (Meta, September 2024)
  • PEFT: LoRA, rank 16, alpha 32, all linear projections
  • Algorithm: GRPO β€” group-standardized advantage with K3 KL anchor
  • Hyperparameters: lr=5e-6, KL Ξ²=0.04, clip Ξ΅=0.2, max episode steps=20

How to use

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "meta-llama/Llama-3.2-3B-Instruct"
adapter = "anshumanatrey/pharmarl-llama-3b-trained-anshuman"

tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, adapter)

Run inside the PharmaRL environment for full episode rollouts against the chemistry oracles.

Links

Submission context

Submitted to the Meta PyTorch OpenEnv Hackathon β€” Round 2 Grand Finale, Bangalore, April 25–26, 2026.

Team AI Mafias β€” Anshuman Atrey, Sahil Shah, Vijay Kota.

Limitations

  • 3B base, hackathon-scale training (200 GRPO steps, single GPU, ~6 hours wall clock).
  • Composite reward is reward-hackable in the sense Renz et al. (2020) document β€” see the molecule audit trail in the run bundle for an example of mid-training reward exploitation at step 100 and recovery at step 150.
  • No JNK3 held-out evaluation completed within the hackathon time budget; the env supports it but the eval pass is pending.
  • This is infrastructure for LLM-as-policy chemistry research, not a clinical or production drug-discovery tool. The model does not produce candidate drugs.

Citation

@misc{atrey2026pharmarl,
  title        = {AI Alchemy in Medicine: A Vision for LLM-as-Policy Molecular Design via {OpenEnv}},
  author       = {Atrey, Anshuman and Shah, Sahil and Kota, Vijay},
  year         = {2026},
  publisher    = {Zenodo},
  url          = {https://zenodo.org/records/19788570},
  note         = {Meta PyTorch OpenEnv Hackathon, Round 2 Grand Finale, Bangalore}
}
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for anshumanatrey/pharmarl-llama-3b-trained-anshuman

Adapter
(857)
this model

Space using anshumanatrey/pharmarl-llama-3b-trained-anshuman 1