Skip to main navigation Skip to search Skip to main content

FLUD: A Scalable and Configurable Systolic Array Design for LU Decomposition on FPGAs

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Lower-upper decomposition (LUD) is one of the most popular matrix factorization techniques in linear algebra and has been widely used in many scientific and engineering applications. While prior studies have investigated various strategies to accelerate block LUD on FPGAs for arbitrary input sizes, they often suffer from one or more of the following limitations: 1) excessive resource utilization due to separate PE (processing element) designs for different matrix blocks with diverse computation patterns; 2) excessive on-chip memory usage due to buffer-based designs; and 3) insufficient parallelism as only one-level parallelism (either row-level or iteration-level parallelism) was exploited due to complex dependencies. To address those limitations, we propose FLUD, a streamingbased systolic array design on the EPGA to accelerate block LUD, which shares the systolic array to accelerate different matrix blocks and exploits both column-level parallelism and iterationlevel parallelism. First, FLUD implements a configurable systolic array that is shared by different LUD blocks and scalable to arbitrary input sizes. To further optimize its hardware resource efficiency, FLUD groups a column of PEs together to replace their FIFO connections with lightweight registers and reduce multiple copies of local control logic inside each PE (for the purpose of resource sharing among different LUD blocks) into a global one. Moreover, FLUD devises a computation schedule to effectively share the highly-optimized systolic array design among the execution of different LUD blocks. Lastly, to enable fast design space exploration on a given FPGA platform, we develop an automation tool to automatically generate the optimized FLUD design in Vitis high-level synthesis (HLS), where users can configure the design size and data precision based on their needs. Experimental results demonstrate that FLUD achieves a peak throughput of 427.95 GFLOPS for single-precision floatingpoint LUD, which is about 3x faster than state-of-the-art FPGA design. Compared to the LAPACK library running on a 12 -core Xeon Silver 4214 CPU, FLUD achieves 4.71x higher throughput and 10.25x better throughput/watt.

Original languageEnglish (US)
Title of host publicationProceedings - International Conference on Field Programmable Technology 2024, ICFPT 2024
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331523213
DOIs
StatePublished - 2024
Externally publishedYes
Event23rd International Conference on Field Programmable Technology, ICFPT 2024 - Sydney, Australia
Duration: Dec 10 2024Dec 12 2024

Publication series

NameProceedings - International Conference on Field-Programmable Technology, ICFPT
ISSN (Print)2837-0430
ISSN (Electronic)2837-0449

Conference

Conference23rd International Conference on Field Programmable Technology, ICFPT 2024
Country/TerritoryAustralia
CitySydney
Period12/10/2412/12/24

Bibliographical note

Publisher Copyright:
© 2024 IEEE.

Fingerprint

Dive into the research topics of 'FLUD: A Scalable and Configurable Systolic Array Design for LU Decomposition on FPGAs'. Together they form a unique fingerprint.

Cite this