AMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs. It holds 16B total parameters but activates only 2.8B per token, using Gated MLA and FarSkip-Collective.
MarkTechPost reports that aMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts language model trained from scratch on Instinct MI300X and MI325X GPUs.
The report adds: It holds 16B total parameters but activates only 2.8B per token, using Gated MLA and FarSkip-Collective.
AMD published weights from every training stage, plus data mixtures, configs, and inference code.
This report is covered across the current DAMMNEWS feeds. Technology stories often mix a product announcement with claims about availability, performance or impact. Those parts should be checked separately.
This is a developing report. DAMMNEWS will surface related coverage as closely matched reporting enters the live cache.
DAMMNEWS RSS BRIEFING — generated locally from the available MarkTechPost headline and RSS excerpt. It is not a summary of the full source article.