Pretraining stack
Full pretraining is rare outside foundation-model teams, but Hugging Face documents compatible stacks.
Adapted for this playbook from the 🤗 Transformers documentation by Hugging Face. Official page: https://huggingface.co/docs/transformers/community_integrations/nanotron. Images © Hugging Face (
documentation-images) unless noted. This is not a substitute for the upstream docs — verify against the current version.Covers Hugging Face pages: community_integrations/nanotron, torchtitan, nemo_automodel_pretraining
| Stack | Doc |
|---|---|
| Nanotron | Nanotron |
| torchtitan | torchtitan |
| NeMo Automodel | NeMo Automodel pretraining |
Most enterprise programmes should fine-tune or RAG rather than pretrain — see Fine-tuning and the playbook RAG guides.
Discussion
Comments​
Share feedback or questions about this page. No account required.
Loading comments…