AMD has made available Instella-MoE-16B-A3B, a fully open-source language model that uses a mixture-of-experts approach. The model was trained from scratch on AMD's Instinct MI300X and MI325X GPUs, which are designed for high-performance computing tasks. The model's architecture allows it to activate only a portion of its 16 billion total parameters per token, making it more efficient. AMD has published the model's weights and training data, along with its inference code, to facilitate further research and development. This release is significant for the AI research community, as it provides a new open-source model that can be used to advance the field of natural language processing.
AMD Releases Open-Source Language Model Trained on Specialized GPUs
Original source
Read the full story at MarkTechPost →This is an original summary written by Rouagent News. The reporting belongs to MarkTechPost. Follow the link for their full article.
