Hey, we still got at least a year or two..
๐ Aniimage model coming soon
Jackson Yeskie
8BitStudio
P(doom) <0.01%
AI & ML interests
Text to image, Language models
Recent Activity
repliedto mithulaartigala's post 3 minutes ago
I am LOVING HuggingFace Pro <3 reacted to mithulaartigala's post with ๐ค about 16 hours ago
I am LOVING HuggingFace Pro <3 liked a model 2 days ago
Abiray/Qwen-Image-2.1-Turbo-GGUFOrganizations
None yet
replied to mithulaartigala's post 3 minutes ago
reacted to mithulaartigala's post with ๐ค about 16 hours ago
upvoted a changelog 6 days ago
Hugging Face Changelog
P(doom) on user profiles
โข 339
WAKE UP
11
#5 opened 7 days ago
by
GGUFGuy
Thats cool! If you can say, how big is the model, im curious
Hugging Science is DEAD???
โ 2
38
#2 opened 16 days ago
by
XBFDI
Message board idea
๐ 5
7
#3 opened 11 days ago
by
el4
Line count
10
#2 opened about 2 months ago
by
Sterling-Ai
reacted to harshitkgupta's post with ๐ 14 days ago
Post
8744
Fine-tuned Qwen 2.5 (0.5B โ 3B) on real coding-agent traces, 10 controlled runs, one 16GB Mac. Compared PyTorch MPS vs. Apple MLX for local LoRA SFT โ and the honest answer is "it depends on what you're optimizing for":
โข PyTorch MPS: 2.2xโ5.7x faster raw throughput, but hits a hard memory wall โ can't load a 3B model in FP16 on 16GB.
โข Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kโ4k tokens).
โข 4-bit quantization doesn't cost you convergence โ eval loss tracks closely across backends.
โข The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it โ dequant only explains 1.07xโ1.4x of the gap. A ~4.1โ4.6x framework-level gap remains either way.
All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed.
Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
โข PyTorch MPS: 2.2xโ5.7x faster raw throughput, but hits a hard memory wall โ can't load a 3B model in FP16 on 16GB.
โข Apple MLX: 4-bit QLoRA fits 3B+ models with almost flat memory scaling as context grows (+109 MB going from 1kโ4k tokens).
โข 4-bit quantization doesn't cost you convergence โ eval loss tracks closely across backends.
โข The bigger surprise: most of MLX's slowdown isn't the 4-bit dequant tax. Two of the 10 runs went unquantized to isolate it โ dequant only explains 1.07xโ1.4x of the gap. A ~4.1โ4.6x framework-level gap remains either way.
All 10 LoRA adapters + Trackio logs are public so the numbers are checkable, not just claimed.
Full writeup: https://huggingface.co/blog/harshitkgupta/fine-tuning-coding-agents-on-mac-pytorch-mps-mlx
replied to Banaxi-Tech's post 16 days ago
does BananaMind have a website? Would love to easily test these on it if so
YOO WHAT?
60
#1 opened 18 days ago
by
Bc-AI
I think they are fine, and they are cool, but is there a clear use for them? What are you for, @BananaMindBot
Saw it being used as a computer automation tool (i.e. open an app, type, something, etc) I think once more open jev type models come out we will have crazy things