$AMD released a state-of-the-art, fully open model trained entirely on MI300X and MI325X GPUs using ROCm


Instella-MoE is a 16-billion-parameter mixture-of-experts model that activates only 2.8 billion parameters per token, reducing inference costs while remaining competitive with dense and MoE models with similar or larger active parameter counts
It supports a 64K context window and was trained through six stages, from pre-training to reinforcement learning
AMD also introduced two new efficiency techniques:
- Gated Multi-head Latent Attention, which filters low-value attention outputs.
- FarSkip-Collective, which overlaps communication and computation, improving pre-training speed by 12.7% and reducing time-to-first-token by up to 39.2%.
AMD-3.29%
post-image
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • 2
  • Repost
  • Share
Comment
Add a comment
Add a comment
GateUser-17ba047d
· 15m ago
How does blockchain technology work
Reply0
Mehedi667
· 21m ago
😉🔥
Reply0
  • Pinned