Skip to content
MangoBoost

MANGO Alphonso™

High EndAMD GPUs

Full Stack Server Optimized for AI

High performance, high efficiency, turnkey AI infrastructure solution.

Turnkey Solution. Managed as one.

Full stack AI system. LLMBoost™, BoostX™ RoCE AI and MAC tune throughput, cost and operations as one.

Full stack
AI system.

Full stack AI system.

Built for large scale training and high volume inference, delivered as one validated turnkey stack.

Network and TCO,
optimized in house.

Network and TCO, optimized in house.

LLMBoost™ and onboard BoostX™ RoCE AI offload the data path to raise throughput and lower total cost of ownership.

Unified system.
One screen.

Unified system. One screen.

MAC (Mango AI Center) manages distributed nodes as a single system, with a real time dashboard for easy operations.

Compute and networking, balanced by design.

Compute and networking balanced to keep every GPU fed, never idle.

MANGO ALPHONSO™ · single node
System Memory
Gen5 x16
CPU 1
Gen5 x16

2× PCIe Gen5 switch

PCIe Switch
PCIe Switch
GPU
MI355X
GPU
MI355X
GPU
MI355X
GPU
MI355X
BoostX™
RoCE AI
BoostX™
RoCE AI
BoostX™
RoCE AI
BoostX™
RoCE AI
Up to 400 Gbps per GPU
Ethernet, RoCE v2, UEC-ready
System Memory
Gen5 x16
CPU 2
Gen5 x16

2× PCIe Gen5 switch

PCIe Switch
PCIe Switch
GPU
MI355X
GPU
MI355X
GPU
MI355X
GPU
MI355X
BoostX™
RoCE AI
BoostX™
RoCE AI
BoostX™
RoCE AI
BoostX™
RoCE AI
Up to 400 Gbps per GPU
Ethernet, RoCE v2, UEC-ready

More performance. Lower cost.

Alphonso™ is tuned as a full stack system, not a collection of parts. That shows up as more throughput from every GPU, and a lower cost to run the cluster over its life.

01 · Performance

Superior performance per dollar.

Same class of GPU, same cluster, same workload. What changes is the server, tuned as a full stack system to deliver more throughput.

Performance/$ = throughput (tok/s/GPU) ÷ modelled 3 year TCO · indexed, baseline = 100 · higher is better

vs MI355X(Base) Server

UP TO
+41.7%

more performance per dollar

100
141.7
MI355X (Base) ServerMango Alphonso

vs NVIDIA B200 Server

UP TO
+47.0%

more performance per dollar

100
147
B200 ServerMango Alphonso

GLM 5/5.1 FP8 · coding agent, 8K in / 1K out · best observed result

02 · TCO

Lower TCO per cluster.

The saving comes from the node and the fabric together, with UEC-ready features on the DPU instead of premium switch silicon.

3 year TCO · 256 GPU cluster · indexed, lower is better

NVIDIA B200 cluster
100
INDEX
Mango Alphonso™
79
INDEX
~21% lower

3 year TCO on a 256 GPU build.

$4.6M saved over 3 years

Modelled 3 year TCO on a 256 GPU build, at matched scale and topology.

Same scale, same topology

256 GPU, 32 nodes, 8 leaf + 4 spine at 1 : 1, 400G end to end.

UEC-ready on the DPU

Packet spraying and congestion control move off the switch, so a standard 400GbE switch is enough.

Ethernet, not a proprietary fabric

RoCE v2 on the spine you already run, with standard optics and standard skills.

Baseline = 8× B200 server · TCO includes power and maintenance · pricing varies by vendor and contract.

Every node, managed as one.

An agent on each node streams live telemetry to a single master, so you monitor, validate and operate the whole cluster from one screen.

Agent AALPHONSO™ node
Agent BALPHONSO™ node
Agent CALPHONSO™ node
Telemetry
Unified control
Mango AI Center · Master
MASTER: RUNNINGSYSTEM HEALTHY
Cluster Overview
Monitoring 3 active nodes
eval2Online
eval4Online
smc21Offline
CPU
42%
Memory
58%
GPU
8/8
Health
OK

Single pane of glass

MAC treats distributed nodes as one system, with a GUI dashboard for every layer of the stack.

Real time health

Cluster overview with live CPU, memory, GPU, temperature and power across every active node.

Package consistency

Version drift checks keep drivers, toolkits and runtimes aligned across every node.

Network and fabric visibility

Per NIC RDMA bandwidth and switch to port topology, visualized in real time.