Data prefetching and data forwarding in shared memory multiprocessors

D. K. Poulsen, Pen Chung Yew

Research output: Contribution to journalConference articlepeer-review

31 Scopus citations


This paper studies and compares the use of data prefetching and an alternative mechanism, data forwarding, for reducing memory latency due to interprocessor communication in cache coherent, shared memory multiprocessors. Two multiprocessor prefetching algorithms are presented and compared. A simple blocked vector prefetching algorithm, considerably less complex than existing software pipelined prefetching algorithms, is shown to be effective in reducing memory latency and increasing performance. A Forwarding Write operation is used to evaluate the effectiveness of forwarding. The use of data forwarding results in significant performance improvements over data prefetching for codes exhibiting less spatial locality. Algorithms for data prefetching and data forwarding are implemented in a parallelizing compiler. Evaluation of the proposed schemes and algorithms is accomplished via execution-driven simulation of large, optimized, parallel numerical application codes with loop-level and vector parallelism. More data, discussion, and experiment details can be found in [1].

Original languageEnglish (US)
Article number5727799
Pages (from-to)II276-II280
JournalProceedings of the International Conference on Parallel Processing
StatePublished - 1994
Event23rd International Conference on Parallel Processing, ICPP 1994 - Raleigh, NC, United States
Duration: Aug 15 1994Aug 19 1994


Dive into the research topics of 'Data prefetching and data forwarding in shared memory multiprocessors'. Together they form a unique fingerprint.

Cite this