Today data scientist is turning to cloud for AI and HPC workloads. However, AI/HPC applications require high computational throughput where generic cloud resources would not suffice. There is a strong demand for OpenStack to support hardware accelerated devices in a dynamic model.
In this session, we will introduce OpenStack Acceleration Service – Cyborg, which provides a management framework for accelerator devices (e.g. FPGA, GPU, NVMe SSD). We will also discuss Rack Scale Design (RSD) technology and explain how physical hardware resources can be dynamically aggregated to meet the AI/HPC requirements. The ability to “compose on the fly” with workload-optimized hardware and accelerator devices through an API allow data center managers to manage these resources in an efficient automated manner.
We will also introduce an enhanced telemetry solution with Gnnochi, bandwidth discovery and smart scheduling, by leveraging RSD technology, for efficient workloads management in HPC/AI cloud.
The attendees will learn
- The background and current state of OpenStack Cyborg project and RSD
- How OpenStack Cyborg manages accelerator devices (e.g. FPGA, GPU, NVMe SSD) with RSD?
- How Cyborg and RSD meet the requirement of HPC/AI?
- Enhanced telemetry solution with Gnnochi for HPC/AI cloud by leveraging RSD technology
- Bandwidth discovery and smart scheduling with RSD API