Presentation
A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC Workloads
DescriptionDataflow accelerators can provide energy-efficient and high-performance alternatives to current popular architectures. However, little work has been done to enable accelerator-initiated, scalable collective communication for these architectures. We develop a high-level synthesis (HLS) interface to bridge this gap through software-hardware co-design. Given the tendency of dataflow applications to use reads and writes to streams to express data transfer, we develop a streaming interface implementing fine-grained transfers to the host processor. Data can then be communicated through MPI and transferred to the receiving accelerator. As a result, the interface uses few hardware resources for communication. As a case study, we enhance the HPL benchmark from HPCC_FPGA with our contributions. We evaluate our final design on up to 16 FPGAs, achieving up to 18% improvement in application throughput and 36% reduced latency in kernel execution. Additionally, we design a stencil benchmark that showcases superlinear speedup in a strong scaling scenario.
Event Type
Paper
TimeThursday, 20 November 20252:15pm - 2:37pm CST
Location275
Programming Frameworks





