Abstract
Lower-upper decomposition (LUD) is one of the most popular matrix factorization techniques in linear algebra and has been widely used in many scientific and engineering applications. While prior studies have investigated various strategies to accelerate block LUD on FPGAs for arbitrary input sizes, they often suffer from one or more of the following limitations: 1) excessive resource utilization due to separate PE (processing element) designs for different matrix blocks with diverse computation patterns; 2) excessive on-chip memory usage due to buffer-based designs; and 3) insufficient parallelism as only one-level parallelism (either row-level or iteration-level parallelism) was exploited due to complex dependencies. To address those limitations, we propose FLUD, a streamingbased systolic array design on the EPGA to accelerate block LUD, which shares the systolic array to accelerate different matrix blocks and exploits both column-level parallelism and iterationlevel parallelism. First, FLUD implements a configurable systolic array that is shared by different LUD blocks and scalable to arbitrary input sizes. To further optimize its hardware resource efficiency, FLUD groups a column of PEs together to replace their FIFO connections with lightweight registers and reduce multiple copies of local control logic inside each PE (for the purpose of resource sharing among different LUD blocks) into a global one. Moreover, FLUD devises a computation schedule to effectively share the highly-optimized systolic array design among the execution of different LUD blocks. Lastly, to enable fast design space exploration on a given FPGA platform, we develop an automation tool to automatically generate the optimized FLUD design in Vitis high-level synthesis (HLS), where users can configure the design size and data precision based on their needs. Experimental results demonstrate that FLUD achieves a peak throughput of 427.95 GFLOPS for single-precision floatingpoint LUD, which is about 3x faster than state-of-the-art FPGA design. Compared to the LAPACK library running on a 12 -core Xeon Silver 4214 CPU, FLUD achieves 4.71x higher throughput and 10.25x better throughput/watt.
| Original language | English (US) |
|---|---|
| Title of host publication | Proceedings - International Conference on Field Programmable Technology 2024, ICFPT 2024 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| ISBN (Electronic) | 9798331523213 |
| DOIs | |
| State | Published - 2024 |
| Externally published | Yes |
| Event | 23rd International Conference on Field Programmable Technology, ICFPT 2024 - Sydney, Australia Duration: Dec 10 2024 → Dec 12 2024 |
Publication series
| Name | Proceedings - International Conference on Field-Programmable Technology, ICFPT |
|---|---|
| ISSN (Print) | 2837-0430 |
| ISSN (Electronic) | 2837-0449 |
Conference
| Conference | 23rd International Conference on Field Programmable Technology, ICFPT 2024 |
|---|---|
| Country/Territory | Australia |
| City | Sydney |
| Period | 12/10/24 → 12/12/24 |
Bibliographical note
Publisher Copyright:© 2024 IEEE.
Fingerprint
Dive into the research topics of 'FLUD: A Scalable and Configurable Systolic Array Design for LU Decomposition on FPGAs'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS