Publications

Publications by Nachiket Kapre and collaborators, in reverse chronological order.

Publication Timeline

Each box links to its publication entry below.

Publication Best Paper Nominee / Runner-up

<input type=”search” id=”bibsearch” spellcheck=”false” autocomplete=”off” class=”search bibsearch-form-input” placeholder=”Type to filter”

2026

  1. A Protocol-Independent Transport Architecture
    Kimiya Mohammadtaheri, David Gao, Samuel Zhang, Matthew Chen, Eric Su, Pengyu Ji, Saad Syed, Chris Neely, Mario Baldi, Nachiket Kapre, and Mina Tahmasbi Arashloo
    May 2026
  2. Jack of All Scales: A Versatile FPGA Tensor Block for MXFP Precisions
    Marwan Mekhemer, Ahmed Elsousy, Balaji Venkatesh, Raphael Rowley, Vaughn Betz, Nachiket Kapre, and Andrew Boutros
    In International Conference on Field-Programmable Logic and Applications, Sep 2026
  3. Cocotb-PYNQ-PR: From Co-Simulation to Deployment, A Unified DFX Framework for PYNQ
    Emir Guevara, Chaitanya Sharma, Gavin Lusby, and Nachiket Kapre
    In International Conference on Field-Programmable Logic and Applications, Sep 2026

2025

  1. Cocotb-Pynq: Co-simulating Python+RTL applications targeting Pynq platforms with Cocotb
    Gavin Lusby and Nachiket Kapre
    In International Conference on Field-Programmable Logic and Applications, Sep 2025

2024

  1. GraphNoC: Graph Neural Networks for Application-Specific FPGA NoC Performance Prediction
    Gurshaant Malik and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2024

2023

  1. Ditty: Directory-based Cache Coherence for Multicore Safety-critical Systems
    Zhuanhao Wu, Marat Bekmyrza, Nachiket Kapre, and Hiren Patel
    In Design, Automation, and Test in Europe, Apr 2023

2022

  1. HopliteML: Evolving application customized FPGA NoCs with adaptable routers and regulators
    Gurshaant Malik, Ian Elmor Lang, Rodolfo Pellizzoni, and Nachiket Kapre
    ACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPL 2020 , 2022
  2. RapidLayout: Fast Hard Block Placement of FPGA-optimized Systolic Arrays using Evolutionary Algorithm
    Niansong Zhang, Xiang Chen, and Nachiket Kapre
    ACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPL 2020 , 2022
  3. Managing HBM Bandwidth on Multi-Die FPGAs with FPGA Overlay NoCs
    Srinirdheeshwar Kuttuva Prakash, Hiren Patel, and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2022

2021

  1. Mocarabe: High-Performance Time-Multiplexed Overlays for FPGAs
    Frederick Tombs, Alireza Mellat, and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2021
  2. Worst-Case Latency Analysis for the Versal NoC Network Packet Switch
    Ian Lang, Nachiket Kapre, and Rodolfo Pellizzoni
    In IEEE/ACM Symposium on Networks-on-Chip, Oct 2021

2020

  1. HopliteBuf: Network Calculus-Based Design of FPGA NoCs with Provably Stall-Free FIFOs
    Tushar Garg, Saud Wasly, Rodolfo Pellizzoni, and Nachiket Kapre
    ACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPGA 2019 , 2020
  2. RapidLayout: Fast Hard Block Placement of FPGA-optimized Systolic Arrays using Evolutionary Algorithms
    Niansong Zhang, Xiang Chen, and Nachiket Kapre
    In International Conference on Field-Programmable Logic and Applications, Sep 2020
  3. Learn the Switches: Evolving FPGA NoCs with Stall-Free and Backpressure based Routers
    Gurshaant Malik, Ian Lang, Rodolfo Pellizzoni, and Nachiket Kapre
    In International Conference on Field-Programmable Logic and Applications, Sep 2020
  4. Exploring the Impact of Switch Arity on Butterfly Fat Tree FPGA NoCs
    Ian Lang, Ziqiang Huang, and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2020

2019

  1. Partitioning FPGA-Optimized Systolic Arrays for Fun and Profit
    Long Chung Chan, Gurshaant Malik, and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2019
  2. Scaling the Cascades: Interconnect-aware FPGA implementation of Machine Learning problems
    Ananda Samajdar, Tushar Garg, Tushar Krishna, and Nachiket Kapre
    In 29th International Conference on Field-Programmable Logic and Applications, Sep 2019
  3. Timing-aware routing in the RapidWright framework
    Leo Liu and Nachiket Kapre
    In 29th International Conference on Field-Programmable Logic and Applications, Sep 2019
  4. Enhancing Butterfly Fat Tree NoCs for FPGAs with lightweight flow control
    Gurshaant Malik and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2019
  5. HopliteBuf: FPGA NoCs with Provably Stall-Free FIFOs
    Tushar Garg, Saud Al Wasly, Rodolfo Pellizzoni, and Nachiket Kapre
    In International Symposium on Field-Programmable Gate Arrays, Feb 2019
  6. RapidRoute: Fast Assembly of Communication Structures for FPGA Overlays
    Leo Liu, Jay Weng, and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2019

2018

  1. Implementing NEF Neural Networks on Embedded FPGAs
    Benjamin Morcos, Terrence Stewart, Chris Eliasmith, and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2018
  2. DaCO: A High-Performance Token Dataflow Coprocessor Overlay for FPGAs
    Siddhartha and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2018
  3. FastTrack: Leveraging Heterogeneous FPGA Wires to Design Low-cost High-performance Soft NoCs
    Nachiket Kapre and Tushar Krishna
    In International Symposium on Computer Architecture, Jun 2018
  4. LegUp-NoC: High-Level Synthesis of Loops with Indirect Addressing
    Asif Islam and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2018
  5. Hoplite-Q: Priority-Aware Routing in FPGA Overlay NoCs
    Siddhartha and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2018

2017

  1. CaffePresso: Accelerating Convolutional Networks on Embedded SoCs
    Gopalakrishna Hegde, Siddhartha, and Nachiket Kapre
    ACM Transactions on Embedded Computing SystemsSpecial Issue: ESWEEK 2016 , 2017
  2. Hoplite: A Deflection-Routed Directional Torus NoC for FPGAs
    Nachiket Kapre and Jan Gray
    ACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPL 2015 , 2017
  3. HopliteRT: An Efficient FPGA NoC for Real-Time Applications
    Saud Al Wasly, Rodolfo Pellizzoni, and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2017
  4. Deflection-Routed Butterfly Fat Trees on FPGAs
    Nachiket Kapre
    In 27th International Conference on Field-Programmable Logic and Applications, Sep 2017
  5. Enabling Partial Reconfiguration and Low Latency Routing using Segmented FPGA NoCs
    Kizhepatt Vipin, Jan Gray, and Nachiket Kapre
    In 27th International Conference on Field-Programmable Logic and Applications, Sep 2017
  6. On Bit-Serial NoCs for FPGAs
    Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2017
  7. Implementing FPGA overlay NoCs using the Xilinx UltraScale memory cascades
    Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2017
  8. eBSP: Managing NoC traffic for BSP workloads on the 16-core Adapteva Epiphany-III Processor
    Siddhartha and Nachiket Kapre
    In Design, Automation, and Test in Europe, Mar 2017
  9. Applying Models of Computation to OpenCL Pipes for FPGA Computing
    Nachiket Kapre and Hiren Patel
    In 5th International Workshop on OpenCL, May 2017
  10. 120-core microAptiv MIPS Overlay for the Terasic DE5-NET FPGA board
    Chethan Kumar H B, Gourav Modi, Prashant Ravi, and Nachiket Kapre
    In International Symposium on Field-Programmable Gate Arrays, Feb 2017

2016

  1. Optimizing Soft Vector Processing in FPGA-based Embedded Systems
    Nachiket Kapre
    ACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPL 2014 , 2016
  2. Deflection Routing for Multi-Level FPGA Overlay NoCs
    Chethan Kumar H B, Shubham Agarwal, and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2016
  3. Preventive Detection of Mosquito Populations using Embedded Machine Learning on Low Power IoT Platforms
    Prashant Ravi, Uma Syam, and Nachiket Kapre
    In Seventh ACM Symposium on Computing and Development, Nov 2016
  4. CaffePresso: An Optimized Library for Deep Learning on Embedded Accelerator-based platforms
    Gopalakrishna Hegde, Siddhartha, Nachiappan Ramasamy, and Nachiket Kapre
    In International Conference on Compilers, Architecture, and Synthesis for Embedded Systems, Oct 2016
  5. Hoplite-DSP: Harnessing the Xilinx DSP48 Multiplexers to efficiently support NoCs on FPGAs
    Chethan Kumar H B and Nachiket Kapre
    In 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
  6. Boosting Convergence of Timing Closure using Feature Selection in a Learning-Driven Approach
    Que Yanghua, Harnhua Ng, and Nachiket Kapre
    In 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
  7. Survey of Domain-Specific Languages for FPGA Computing
    Nachiket Kapre and Samuel Bayliss
    In 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
  8. Marathon: Statically-Scheduled Conflict-Free Routing on FPGA Overlay NoCs
    Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2016
  9. GPU-Accelerated High-Level Synthesis for Bitwidth Optimization of FPGA Datapaths
    Nachiket Kapre and Ye Deheng
    In International Symposium on Field-Programmable Gate Arrays, Feb 2016
  10. Vector FPGA Acceleration of 1-D DWT Computations using Sparse Matrix Skeletons
    Sidharth Maheshwari, Gourav Modi, Siddhartha, and Nachiket Kapre
    In 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
  11. Improving Classification Accuracy of a Machine Learning approach for FPGA Timing Closure
    Que Yanghua, Nachiket Kapre, Harnhua Ng, and Kirvy Teo
    In International Symposium on Field-Programmable Custom Computing Machines, May 2016
  12. Case for Design-Specific Machine Learning in Timing Closure of FPGA Designs
    Que Yanghua, Chinnakkannu Adaikkal Raj, Harnhua Ng, Kirvy Teo, and Nachiket Kapre
    In International Symposium on Field-Programmable Gate Arrays, Feb 2016
  13. Evaluating Embedded FPGA Accelerators for Deep Learning Applications
    Gopalakrishna Hegde, Siddhartha, Nachiappan Ramasamy, Vamsi Buddha, and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2016
  14. Communication Optimization for the 16-core Epiphany Floating-Point Processor Array
    Siddhartha and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2016
  15. Machine-Learning driven Auto-Tuning of High-Level Synthesis for FPGAs
    Li Ting, Harri Sapto Wijaya, and Nachiket Kapre
    In International Symposium on Field-Programmable Gate Arrays, Feb 2016

2015

  1. A Case for Embedded FPGA-based SoCs in Energy-Efficient Acceleration of Graph Problems
    Pradeep Moorthy and Nachiket Kapre
    Supercomputing Frontiers and InnovationsSpecial Best Papers Issue from Supercomputing Frontiers 2015 , 2015
  2. Communication Optimization of Iterative Sparse Matrix-Vector Multiply on GPUs and FPGAs
    Abid Rafique, George Constantinides, and Nachiket Kapre
    IEEE Transactions on Parallel and Distributed Systems, Jan 2015
  3. Hoplite: Building Austere Overlay NoCs for FPGAs
    Nachiket Kapre and Jan Gray
    In 25th International Conference on Field-Programmable Logic and Applications, Sep 2015
  4. Limits of FPGA Acceleration of 3D Green’s Function Computation for Geophysical Applications
    Nachiket Kapre, Selvakumar Jayakrishnan, Parjanya Gupta, Sagar Masuti, and Sylvain Barbot
    In 25th International Conference on Field-Programmable Logic and Applications, Sep 2015
  5. Custom FPGA-based Soft-Processors for Sparse Graph Acceleration
    Nachiket Kapre
    In 26th IEEE International Conference on Application-specific Systems, Architectures and Processors, Jul 2015
  6. GraphMMU: Memory Management Unit for Sparse Graph Accelerators
    Nachiket Kapre, Han Jianglei, Andrew Bean, Pradeep Moorthy, and Siddhartha
    In 22nd Reconfigurable Architectures WorkshopCo-located with IPDPS 2015 , May 2015
  7. Enhancing Speedups for FPGA Accelerated SPICE through Frequency Scaling and Precision Reduction
    Lim Hui Hui and Nachiket Kapre
    In 22nd Reconfigurable Architectures WorkshopCo-located with IPDPS 2015 , May 2015
  8. Zedwulf: Power-Performance Tradeoffs of a 32-node Zynq SoC cluster
    Pradeep Moorthy and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2015
  9. Driving Timing Convergence of FPGA Designs through Machine Learning and Cloud Computing
    Nachiket Kapre, Bibin Chandrashekaran, Harnhua Ng, and Kirvy Teo
    In International Symposium on Field-Programmable Custom Computing Machines, May 2015
  10. Energy-Efficient Acceleration of OpenCV Saliency Computation using Soft Vector Processors
    Gopalakrishna Hegde and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2015
  11. On Data Forwarding in Deeply Pipelined Soft Processor
    Hui Yan Cheah, Suhaib A. Fahmy, and Nachiket Kapre
    In International Symposium on Field-Programmable Gate Arrays, Feb 2015
  12. InTime: A Machine Learning Approach for Efficient Selection of FPGA CAD Tool Parameters
    Nachiket Kapre, Harnhua Ng, Kirvy Teo, and Jaco Naude
    In International Symposium on Field-Programmable Gate Arrays, Feb 2015
  13. Sparse Graph Processing using Soft-Processors
    Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2015
  14. FPGA Acceleration of Irregular Iterative Computations using Criticality-Aware Dataflow Optimizations
    Siddhartha and Nachiket Kapre
    In International Symposium on Field-Programmable Gate Arrays, Feb 2015

2014

  1. Relax-Miracle: GPU Parallelization of Semi-Analytic Fourier-Domain solvers for Earthquake Modeling
    Sagar Masuti, Sylvain Barbot, and Nachiket Kapre
    In International Conference on High Performance Computing, Dec 2014
  2. Comparing Soft and Hard Vector Processing in FPGA-based Embedded Systems
    Soh Jun Jie and Nachiket Kapre
    In International Conference on Field-Programmable Logic and Applications, Sep 2014
  3. Limits of Statically Scheduled Token Dataflow Processing
    Nachiket Kapre and Siddhartha
    In 4th International Workshop on Data-Flow Execution Models for Extreme Scale ComputingCo-located with PACT 2014 , Aug 2014
  4. Fanout Decomposition Dataflow Optimizations for FPGA-based Sparse LU Factorization
    Siddhartha and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2014
  5. Analysis and Optimization of a Deeply Pipelined FPGA Soft Processor
    Hui Yan Cheah, Suhaib A. Fahmy, and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2014
  6. Heterogeneous Dataflow Architectures for FPGA-based Sparse LU Factorization
    Siddhartha and Nachiket Kapre
    In International Conference on Field-Programmable Logic and Applications, Sep 2014
  7. Breaking Sequential Dependencies in FPGA-based Sparse LU Factorization
    Siddhartha and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2014
  8. MixFX-SCORE: Heterogeneous Fixed-Point Compilation of Dataflow Computations
    Ye Deheng and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2014
  9. Timing Fault Detection in FPGA-based Circuits
    Edward Stott, Joshua M. Levine, Peter Y. K. Cheung, and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, May 2014
  10. Measuring Timing Errors in FPGA-based Circuits
    Joshua Levine, Edward Stott, and Nachiket Kapre
    In The 10th IEEE Workshop on Silicon Errors in Logic - System Effects, Apr 2014

2013

  1. System-Level FPGA Device Driver with High-Level Synthesis Support
    Kizheppatt Vipin, Shanker Shreejith, Dulitha Gunasekera, Suhaib A Fahmy, and Nachiket Kapre
    In International Conference on Field-Programmable Technology, Dec 2013
  2. Exploiting Input Parameter Uncertainty for Reducing Datapath Precision of SPICE Device Models
    Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2013
  3. Application Composition and Communication Optimization in Iterative Solvers using FPGAs
    Abid Rafique, Nachiket Kapre, and George Constantinides
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2013
  4. Accelerating the SPICE Circuit Simulator using an FPGA - A Case Study
    Nachiket Kapre and André DeHon
    In High-Performance Computing using FPGAs, 2013

2012

  1. SPICE²: Spatial Processors Interconnected for Concurrent Execution for accelerating the SPICE Circuit Simulator using an FPGA
    Nachiket Kapre and André DeHon
    Transactions in CADSpecial Issue on Parallel CAD; volume 31, issue 1 , Jan 2012
  2. Enhancing Performance of Tall-Skinny QR factorization using FPGAs
    Abid Rafique, Nachiket Kapre, and George Constantinides
    In International Conference on Field-Programmable Logic and Applications, Aug 2012
  3. FX-SCORE: A Framework for Fixed-Point Compilation of SPICE Device Models using Gappa++
    Hélène Martorell and Nachiket Kapre
    In International Symposium on Field-Programmable Custom Computing Machines, Apr 2012
  4. A High Throughput FPGA-based Implementation of the Lanczos Method for the Symmetric Extremal Eigenvalue Problem
    Abid Rafique, Nachiket Kapre, and George Constantinides
    In International Symposium on Applied Reconfigurable Computing, Mar 2012

2011

  1. Spatial Hardware Implementation for Sparse Graph Algorithms in GraphStep
    Michael deLorimier, Nachiket Kapre, Nikil Mehta, and André DeHon
    ACM Transactions on Autonomous and Adaptive SystemsSpatial Computing Special Issue , Sep 2011
  2. An NoC Traffic Compiler for efficient FPGA implementation of Sparse Graph-Oriented Workloads
    Nachiket Kapre and André DeHon
    International Journal of Reconfigurable ComputingVolume 2011, Article ID 745147 , 2011
  3. VLIW-SCORE: Beyond C for Sequential Control of SPICE FPGA Acceleration
    Nachiket Kapre and André DeHon
    In International Conference on Field-Programmable Technology, Dec 2011

2010

  1. SPICE²: Spatial Processors Interconnected for Concurrent Execution for accelerating the SPICE Circuit Simulator using an FPGA
    Nachiket Kapre and André DeHon
    In The First Workshop on the Intersections of Computer Architecture and Reconfigurable Logic, Dec 2010
  2. An NoC Traffic Compiler for efficient FPGA implementation of Parallel Graph Applications
    Nachiket Kapre and André DeHon
    In Reconfigurable Communication-centric Systems on Chip, May 2010

2009

  1. Pipelining Saturated Accumulation
    Karl Papadantonakis, Nachiket Kapre, Stephanie Chan, and André DeHon
    IEEE Transactions on Computers, Feb 2009
  2. Parallelizing Sparse Matrix-Solve for SPICE Circuit Simulation using FPGAs
    Nachiket Kapre and André DeHon
    In International Conference on Field-Programmable Technology, Dec 2009
  3. Performance Comparison of Single-Precision SPICE Model-Evaluation on FPGA, GPU, Cell, and Multi-Core Processors
    Nachiket Kapre and André DeHon
    In International Conference on Field-Programmable Logic and Applications, Sep 2009
  4. Accelerating SPICE Model-Evaluation using FPGAs
    Nachiket Kapre and André DeHon
    In IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2009

2008

  1. Programming FPGA Applications in VHDL
    Nachiket Kapre and André DeHon
    In Reconfigurable Computing: The Theory and Practice of FPGA-based Computation, 2008

2007

  1. Optimistic Parallelization of Floating-Point Accumulation
    Nachiket Kapre and André DeHon
    In IEEE Symposium on Computer Arithmetic, Jun 2007

2006

  1. Packet-Switched vs. Time-Multiplexed FPGA Overlay Networks
    Nachiket Kapre, Nikil Mehta, Michael deLorimier, Raphael Rubin, Henry Barnor, Michael Wilson, Michael Wrighton, and André DeHon
    In IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2006
  2. GraphStep: A System Architecture for Sparse Graph Algorithms
    Michael deLorimier, Nachiket Kapre, Nikil Mehta, Dominic Rizzo, Ian Eslick, Raphael Rubin, Tomas Uribe, Thomas Knight Jr., and André DeHon
    In IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2006

2005

  1. Pipelining Saturated Accumulation
    Karl Papadantonakis, Nachiket Kapre, Stephanie Chan, and André DeHon
    In International Conference on Field-Programmable Technology, Dec 2005

2004

  1. Design Patterns for Reconfigurable Computing
    André DeHon, Joshua Adams, Michael deLorimier, Nachiket Kapre, Yuki Matsuda, Helia Naeimi, Michael Vanier, and Michael Wrighton
    In IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2004
  2. Saliency on a chip: a digital approach with an FPGA
    Nachiket Kapre, Dirk Walther, Christof Koch, and André DeHon
    The Neuromorphic EngineerVolume 1, issue 2, Autumn 2004 , 2004