Question
A team is designing training for a large language model that cannot fit on one GPU, and the corpus is large enough to require high throughput. The cluster has fast links inside each GPU node and slower links between nodes. Which strategy is most defensible?