research-article

FORMLESS: scalable utilization of embedded manycores in streaming applications

Authors:
Matin Hashemi

Sharif University of Technology

Sharif University of Technology
View Profile

,
Mohammad H. Foroozannejad

University of California, Davis

University of California, Davis
View Profile

,
Soheil Ghiasi

University of California, Davis

University of California, Davis
View Profile

,
Christoph Etzel

University of Augsburg

University of Augsburg
View Profile

LCTES '12: Proceedings of the 13th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, Tools and Theory for Embedded SystemsJune 2012Pages 71–78https://doi.org/10.1145/2248418.2248429

Published:12 June 2012Publication History

LCTES '12: Proceedings of the 13th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, Tools and Theory for Embedded Systems

Pages 71–78

ABSTRACT

Variants of dataflow specification models are widely used to synthesize streaming applications for distributed-memory parallel processors. We argue that current practice of specifying streaming applications using rigid dataflow models, implicitly prohibits a number of platform oriented optimizations and hence limits portability and scalability with respect to number of processors. We motivate Functionally-cOnsistent stRucturally-MalLEabe Streaming Specification, dubbed FORMLESS, which refers to raising the abstraction level beyond fixed-structure dataflow to address its portability and scalability limitations. To demonstrate the potential of the idea, we develop a design space exploration scheme to customize the application specification to better fit the target platform. Experiments with several common streaming case studies demonstrate improved portability and scalability over conventional dataflow specification models, and confirm the effectiveness of our approach.

References

S. Battacharyya, E. Lee, and P. Murthy. Software synthesis from dataflow graphs. Kluwer Academic Publishers, 1996. Google ScholarDigital Library
S. Stuijk, M. Geilen, and T. Basten. Throughput-buffering trade-off exploration for cyclo-static and synchronous dataflow graphs. IEEE Transactions on Computers, 57(10):1331--1345, 2008. Google ScholarDigital Library
M. Gordon. Compiler techniques for scalable performance of stream programs on multicore architectures. PhD thesis, Massachusetts Institute of Technology, 2010. Google ScholarDigital Library
Andy D. Pimentel et al. Exploring embedded-systems architectures with Artemis. IEEE Computer, 34(11):57--63, 2001. Google ScholarDigital Library
A. Sangiovanni-Vincentelli et al. Benefits and challenges for platform-based design. Design Automation Conference (DAC), pages 409--414, 2004. Google ScholarDigital Library
D. Truong et al. A 167--processor 65 nm computational platform with per-processor dynamic supply voltage and dynamic clock frequency scaling. Symposium on VLSI Circuits, 2008.Google Scholar
S. Bell et al. TILE64 processor: A 64-core SoC with mesh interconnect. International Solid-State Circuits Conference (ISSCC), 2008.Google Scholar
E. Lee and D. Messerschmitt. Synchronous data flow. Proceedings of the IEEE, 75(9):1235--1245, 1987.Google ScholarCross Ref
M. Geilen and T. Basten. Reactive process networks. International Conference on Embedded Software (EMSOFT), pages 137--146, 2004. Google ScholarDigital Library
J. Colaço, A. Girault, G. Hamon, and M. Pouzet. Towards a higher-order synchronous data-flow language. International Conference on Embedded Software (EMSOFT), pages 230--239, 2004. Google ScholarDigital Library
W. Taha. A gentle introduction to multi-stage programming. Domain-Specific Program Generation, Lecture Notes in Computer Science (LNCS), 2004.Google Scholar
J. Adam Cataldo. The power of higher-order composition languages in system design. PhD thesis, University of California, Berkeley, 2006.Google Scholar
Marc Geilen. Reduction techniques for synchronous dataflow graphs. Design Automation Conference (DAC), 2009. Google ScholarDigital Library
B. Bhattacharya and S. Bhattacharyya. Parameterized dataflow modeling for DSP systems. IEEE Transactions on Signal Processing, 49(10):2408--2421, 2001. Google ScholarDigital Library
B.D. Theelen et al. A scenario-aware data flow model for combined long-run average and worst-case performance analysis. Formal Methods and Models in CoDesign, 2006.Google ScholarDigital Library
Maarten H. Wiggers, Marco J. G. Bekooij, and Gerard J. M. Smit. Buffer capacity computation for throughput constrained streaming applications with data-dependent inter-task communication. IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS), 2008. Google ScholarDigital Library
Pascal Fradet, Alain Girault, and Peter Poplavko. A schedulable parametric data-flow MoC. Design, Automation, and Test in Europe (DATE), 2012.Google Scholar
J. Nickolls et al. Scalable parallel programming with CUDA. ACM Queue, 6:40--53, March 2008. Google ScholarDigital Library
CUDA C best practices guide, chapter 4.4. March 2011.Google Scholar
G. Karypis and V. Kumar. METIS 4.0: Unstructured graph partitioning and sparse matrix ordering system. Technical report, CS Dept., University of Minnesota, Minneapolis, 1998.Google Scholar
T. Mohsenin, D. Truong, and B. Baas. Multi-split-row threshold decoding implementations for LDPC codes. International Symposium on Circuits and Systems (ISCAS), 2009.Google ScholarCross Ref
Po-Kuan Huang, Matin Hashemi, and Soheil Ghiasi. System-level performance estimation for application-specific mpsoc interconnect synthesis. Symposium on Application Specific Processors (SASP), 2008. Google ScholarDigital Library
Matin Hashemi. Automated Software Synthesis for Streaming Applications on Embedded Manycore Processors. PhD thesis, University of California, Davis, 2011. Chapter 4. Google ScholarDigital Library

Recommendations

FORMLESS: scalable utilization of embedded manycores in streaming applications
LCTES '12

Variants of dataflow specification models are widely used to synthesize streaming applications for distributed-memory parallel processors. We argue that current practice of specifying streaming applications using rigid dataflow models, implicitly ...
Read More
Multidimensional DSP Core Synthesis for FPGA

Current rapid synthesis approaches for reusable dedicated hardware components (cores) for digital signal processing systems are ineffective since they fail to capture and exploit the manner in which the resulting components are used as part of a ...
Read More
Multidimensional Dataflow Graph Modeling and Mapping for Efficient GPU Implementation
SIPS '12: Proceedings of the 2012 IEEE Workshop on Signal Processing Systems

Multidimensional synchronous dataflow (MDSDF) provides an effective model of computation for a variety of multidimensional DSP systems that have static dataflow structures. In this paper, we develop new methods for optimized implementation of MDSDF ...
Read More

Comments

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Publication

Published in
LCTES '12: Proceedings of the 13th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, Tools and Theory for Embedded Systems
June 2012
153 pages
ISBN:9781450312127
DOI:10.1145/2248418
General Chair:
Reinhard Wilhelm
Saarland University Saarbrücken
,
Program Chairs:
Heiko Falk
Ulm University
,
Wang Yi
Uppsala University
ACM SIGPLAN Notices Volume 47, Issue 5
LCTES '12
MAY 2012
152 pages
ISSN:0362-1340
EISSN:1558-1160
DOI:10.1145/2345141
Issue’s Table of Contents
Copyright © 2012 ACM
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]
Sponsors
In-Cooperation
Publisher
Association for Computing Machinery
New York, NY, United States
Publication History
- Published: 12 June 2012
Permissions
Request permissions about this article.
Request Permissions

Check for updates
Author Tags
dataflow graph
embedded many-core processor
stream application
Qualifiers
- research-article
Conference

Acceptance Rates
Overall Acceptance Rate116of438submissions,26%
Funding Sources
Other Metrics
View Article Metrics

Article Metrics
- 6
  Total Citations
  View Citations
- 223
  Total Downloads
- Downloads (Last 12 months)0
- Downloads (Last 6 weeks)0
Other Metrics
View Author Metrics
Cited By
View all

PDF Format

View or Download as a PDF file.

PDF

eReader

View online with eReader.

eReader

FORMLESS: scalable utilization of embedded manycores in streaming applications

LCTES '12: Proceedings of the 13th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, Tools and Theory for Embedded Systems

ABSTRACT

References

Cited By

Recommendations

FORMLESS: scalable utilization of embedded manycores in streaming applications

Multidimensional DSP Core Synthesis for FPGA

Multidimensional Dataflow Graph Modeling and Mapping for Efficient GPU Implementation