On the AI side, we learn that Meta’s Llama 3 training has not been a smooth ride. In fact, the H100 was the cause of numerous crashes, partly due to faulty memory. It should also be noted that training for this AI lasted 54 days.
NVIDIA’s H100s gave Meta a hard time!
For the record, Llama 3 was trained with an astronomical number of graphics cards. We’re talking here about a cluster of 16,384, all NVIDIA H100s, the most powerful card currently available in this sector.
As a reminder, we’re talking about a graphics card equipped with a GH100 GPU and 80 GB of HBM3 memory. As for the GPU, depending on the variant used, we’re talking about 114 or 132 SM, i.e. a cuda core count of 14,592 or 16,896, depending on whether we’re talking about a PCIe or SXM5 card.
In short, during this 54-day training session, there were a great many problems. In fact, we’re talking about almost a thousand problems. Our colleagues report:
- 419 unexpected failures
- 47 planned maintenance interruptions
- 466 breakdowns
This leaves us with a total of 885 hardware-related errors, broken down as follows: 30.1% NVLink-related and 17.2% HBM3 memory-related. Finally, this leaves room for only two CPU-related errors… Two errors in 54 days of training, that’s crazy!
![[LAB]: AGI TURBOJET UD858 DDR5 RGB 6000 MT/s CL36 AGI TURBOJET UD858 miniature](https://en.overclocking.com/wp-content/medias/sites/4/2026/09/AGI-DDR5-UD858_main-218x150.jpg)






![[Modding] Cooler Master MasterFrame 360 Stage LCD MasterFrame 360 STAGE LCD](https://en.overclocking.com/wp-content/medias/sites/4/2026/06/MasterFrame-360-STAGE-LCD_11-218x150.jpg)



![[Tweak League] PC build: My Checklist & tips](https://en.overclocking.com/wp-content/medias/sites/4/2024/07/overclocking-checlist-montage-pc-218x150.png)
