Publications
Publications by Nachiket Kapre and collaborators, in reverse chronological order.
Publication Timeline
Each box links to its publication entry below.
Publication Best Paper Nominee / Runner-up
<input type=”search” id=”bibsearch” spellcheck=”false” autocomplete=”off” class=”search bibsearch-form-input” placeholder=”Type to filter”
2026
-
- Cocotb-PYNQ-PR: From Co-Simulation to Deployment, A Unified DFX Framework for PYNQIn International Conference on Field-Programmable Logic and Applications, Sep 2026
2025
2024
- GraphNoC: Graph Neural Networks for Application-Specific FPGA NoC Performance PredictionIn International Conference on Field-Programmable Technology, Dec 2024
2023
- Ditty: Directory-based Cache Coherence for Multicore Safety-critical SystemsIn Design, Automation, and Test in Europe, Apr 2023
2022
- RapidLayout: Fast Hard Block Placement of FPGA-optimized Systolic Arrays using Evolutionary AlgorithmACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPL 2020 , 2022
- Managing HBM Bandwidth on Multi-Die FPGAs with FPGA Overlay NoCsIn International Symposium on Field-Programmable Custom Computing Machines, May 2022
2021
- Worst-Case Latency Analysis for the Versal NoC Network Packet SwitchIn IEEE/ACM Symposium on Networks-on-Chip, Oct 2021
2020
- RapidLayout: Fast Hard Block Placement of FPGA-optimized Systolic Arrays using Evolutionary AlgorithmsIn International Conference on Field-Programmable Logic and Applications, Sep 2020
- Learn the Switches: Evolving FPGA NoCs with Stall-Free and Backpressure based RoutersIn International Conference on Field-Programmable Logic and Applications, Sep 2020
2019
- Partitioning FPGA-Optimized Systolic Arrays for Fun and ProfitIn International Conference on Field-Programmable Technology, Dec 2019
- HopliteBuf: FPGA NoCs with Provably Stall-Free FIFOsIn International Symposium on Field-Programmable Gate Arrays, Feb 2019
- RapidRoute: Fast Assembly of Communication Structures for FPGA OverlaysIn International Symposium on Field-Programmable Custom Computing Machines, Apr 2019
2018
- Implementing NEF Neural Networks on Embedded FPGAsIn International Conference on Field-Programmable Technology, Dec 2018
- DaCO: A High-Performance Token Dataflow Coprocessor Overlay for FPGAsIn International Conference on Field-Programmable Technology, Dec 2018
- FastTrack: Leveraging Heterogeneous FPGA Wires to Design Low-cost High-performance Soft NoCsIn International Symposium on Computer Architecture, Jun 2018
- LegUp-NoC: High-Level Synthesis of Loops with Indirect AddressingIn International Symposium on Field-Programmable Custom Computing Machines, Apr 2018
- Hoplite-Q: Priority-Aware Routing in FPGA Overlay NoCsIn International Symposium on Field-Programmable Custom Computing Machines, Apr 2018
2017
- Hoplite: A Deflection-Routed Directional Torus NoC for FPGAsACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPL 2015 , 2017
- Enabling Partial Reconfiguration and Low Latency Routing using Segmented FPGA NoCsIn 27th International Conference on Field-Programmable Logic and Applications, Sep 2017
- On Bit-Serial NoCs for FPGAsIn International Symposium on Field-Programmable Custom Computing Machines, May 2017
- Implementing FPGA overlay NoCs using the Xilinx UltraScale memory cascadesIn International Symposium on Field-Programmable Custom Computing Machines, May 2017
- Applying Models of Computation to OpenCL Pipes for FPGA ComputingIn 5th International Workshop on OpenCL, May 2017
- 120-core microAptiv MIPS Overlay for the Terasic DE5-NET FPGA boardIn International Symposium on Field-Programmable Gate Arrays, Feb 2017
2016
- Optimizing Soft Vector Processing in FPGA-based Embedded SystemsACM Transactions on Reconfigurable Technology and SystemsSpecial Issue: FPL 2014 , 2016
- Deflection Routing for Multi-Level FPGA Overlay NoCsIn International Conference on Field-Programmable Technology, Dec 2016
- Preventive Detection of Mosquito Populations using Embedded Machine Learning on Low Power IoT PlatformsIn Seventh ACM Symposium on Computing and Development, Nov 2016
- CaffePresso: An Optimized Library for Deep Learning on Embedded Accelerator-based platformsIn International Conference on Compilers, Architecture, and Synthesis for Embedded Systems, Oct 2016
- Hoplite-DSP: Harnessing the Xilinx DSP48 Multiplexers to efficiently support NoCs on FPGAsIn 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
- Boosting Convergence of Timing Closure using Feature Selection in a Learning-Driven ApproachIn 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
- Survey of Domain-Specific Languages for FPGA ComputingIn 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
- Marathon: Statically-Scheduled Conflict-Free Routing on FPGA Overlay NoCsIn International Symposium on Field-Programmable Custom Computing Machines, May 2016
- GPU-Accelerated High-Level Synthesis for Bitwidth Optimization of FPGA DatapathsIn International Symposium on Field-Programmable Gate Arrays, Feb 2016
- Vector FPGA Acceleration of 1-D DWT Computations using Sparse Matrix SkeletonsIn 26th International Conference on Field-Programmable Logic and Applications, Sep 2016
- Improving Classification Accuracy of a Machine Learning approach for FPGA Timing ClosureIn International Symposium on Field-Programmable Custom Computing Machines, May 2016
- Case for Design-Specific Machine Learning in Timing Closure of FPGA DesignsIn International Symposium on Field-Programmable Gate Arrays, Feb 2016
- Evaluating Embedded FPGA Accelerators for Deep Learning ApplicationsIn International Symposium on Field-Programmable Custom Computing Machines, May 2016
- Communication Optimization for the 16-core Epiphany Floating-Point Processor ArrayIn International Symposium on Field-Programmable Custom Computing Machines, May 2016
- Machine-Learning driven Auto-Tuning of High-Level Synthesis for FPGAsIn International Symposium on Field-Programmable Gate Arrays, Feb 2016
2015
- A Case for Embedded FPGA-based SoCs in Energy-Efficient Acceleration of Graph ProblemsSupercomputing Frontiers and InnovationsSpecial Best Papers Issue from Supercomputing Frontiers 2015 , 2015
- Communication Optimization of Iterative Sparse Matrix-Vector Multiply on GPUs and FPGAsIEEE Transactions on Parallel and Distributed Systems, Jan 2015
- Hoplite: Building Austere Overlay NoCs for FPGAsIn 25th International Conference on Field-Programmable Logic and Applications, Sep 2015
- Limits of FPGA Acceleration of 3D Green’s Function Computation for Geophysical ApplicationsIn 25th International Conference on Field-Programmable Logic and Applications, Sep 2015
- Custom FPGA-based Soft-Processors for Sparse Graph AccelerationIn 26th IEEE International Conference on Application-specific Systems, Architectures and Processors, Jul 2015
- GraphMMU: Memory Management Unit for Sparse Graph AcceleratorsIn 22nd Reconfigurable Architectures WorkshopCo-located with IPDPS 2015 , May 2015
- Enhancing Speedups for FPGA Accelerated SPICE through Frequency Scaling and Precision ReductionIn 22nd Reconfigurable Architectures WorkshopCo-located with IPDPS 2015 , May 2015
- Zedwulf: Power-Performance Tradeoffs of a 32-node Zynq SoC clusterIn International Symposium on Field-Programmable Custom Computing Machines, May 2015
- Driving Timing Convergence of FPGA Designs through Machine Learning and Cloud ComputingIn International Symposium on Field-Programmable Custom Computing Machines, May 2015
- Energy-Efficient Acceleration of OpenCV Saliency Computation using Soft Vector ProcessorsIn International Symposium on Field-Programmable Custom Computing Machines, May 2015
- On Data Forwarding in Deeply Pipelined Soft ProcessorIn International Symposium on Field-Programmable Gate Arrays, Feb 2015
- InTime: A Machine Learning Approach for Efficient Selection of FPGA CAD Tool ParametersIn International Symposium on Field-Programmable Gate Arrays, Feb 2015
- Sparse Graph Processing using Soft-ProcessorsIn International Symposium on Field-Programmable Custom Computing Machines, May 2015
- FPGA Acceleration of Irregular Iterative Computations using Criticality-Aware Dataflow OptimizationsIn International Symposium on Field-Programmable Gate Arrays, Feb 2015
2014
- Relax-Miracle: GPU Parallelization of Semi-Analytic Fourier-Domain solvers for Earthquake ModelingIn International Conference on High Performance Computing, Dec 2014
- Comparing Soft and Hard Vector Processing in FPGA-based Embedded SystemsIn International Conference on Field-Programmable Logic and Applications, Sep 2014
- Limits of Statically Scheduled Token Dataflow ProcessingIn 4th International Workshop on Data-Flow Execution Models for Extreme Scale ComputingCo-located with PACT 2014 , Aug 2014
- Fanout Decomposition Dataflow Optimizations for FPGA-based Sparse LU FactorizationIn International Conference on Field-Programmable Technology, Dec 2014
- Analysis and Optimization of a Deeply Pipelined FPGA Soft ProcessorIn International Conference on Field-Programmable Technology, Dec 2014
- Heterogeneous Dataflow Architectures for FPGA-based Sparse LU FactorizationIn International Conference on Field-Programmable Logic and Applications, Sep 2014
- Breaking Sequential Dependencies in FPGA-based Sparse LU FactorizationIn International Symposium on Field-Programmable Custom Computing Machines, May 2014
- MixFX-SCORE: Heterogeneous Fixed-Point Compilation of Dataflow ComputationsIn International Symposium on Field-Programmable Custom Computing Machines, May 2014
- Timing Fault Detection in FPGA-based CircuitsIn International Symposium on Field-Programmable Custom Computing Machines, May 2014
- Measuring Timing Errors in FPGA-based CircuitsIn The 10th IEEE Workshop on Silicon Errors in Logic - System Effects, Apr 2014
2013
- Application Composition and Communication Optimization in Iterative Solvers using FPGAsIn International Symposium on Field-Programmable Custom Computing Machines, Apr 2013
- Accelerating the SPICE Circuit Simulator using an FPGA - A Case StudyIn High-Performance Computing using FPGAs, 2013
2012
- SPICE²: Spatial Processors Interconnected for Concurrent Execution for accelerating the SPICE Circuit Simulator using an FPGATransactions in CADSpecial Issue on Parallel CAD; volume 31, issue 1 , Jan 2012
- Enhancing Performance of Tall-Skinny QR factorization using FPGAsIn International Conference on Field-Programmable Logic and Applications, Aug 2012
- A High Throughput FPGA-based Implementation of the Lanczos Method for the Symmetric Extremal Eigenvalue ProblemIn International Symposium on Applied Reconfigurable Computing, Mar 2012
2011
- Spatial Hardware Implementation for Sparse Graph Algorithms in GraphStepACM Transactions on Autonomous and Adaptive SystemsSpatial Computing Special Issue , Sep 2011
- An NoC Traffic Compiler for efficient FPGA implementation of Sparse Graph-Oriented WorkloadsInternational Journal of Reconfigurable ComputingVolume 2011, Article ID 745147 , 2011
- VLIW-SCORE: Beyond C for Sequential Control of SPICE FPGA AccelerationIn International Conference on Field-Programmable Technology, Dec 2011
2010
- SPICE²: Spatial Processors Interconnected for Concurrent Execution for accelerating the SPICE Circuit Simulator using an FPGAIn The First Workshop on the Intersections of Computer Architecture and Reconfigurable Logic, Dec 2010
- An NoC Traffic Compiler for efficient FPGA implementation of Parallel Graph ApplicationsIn Reconfigurable Communication-centric Systems on Chip, May 2010
2009
-
- Parallelizing Sparse Matrix-Solve for SPICE Circuit Simulation using FPGAsIn International Conference on Field-Programmable Technology, Dec 2009
- Performance Comparison of Single-Precision SPICE Model-Evaluation on FPGA, GPU, Cell, and Multi-Core ProcessorsIn International Conference on Field-Programmable Logic and Applications, Sep 2009
- Accelerating SPICE Model-Evaluation using FPGAsIn IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2009
2008
- Programming FPGA Applications in VHDLIn Reconfigurable Computing: The Theory and Practice of FPGA-based Computation, 2008
2007
- Optimistic Parallelization of Floating-Point AccumulationIn IEEE Symposium on Computer Arithmetic, Jun 2007
2006
- Packet-Switched vs. Time-Multiplexed FPGA Overlay NetworksIn IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2006
- GraphStep: A System Architecture for Sparse Graph AlgorithmsIn IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2006
2005
- Pipelining Saturated AccumulationIn International Conference on Field-Programmable Technology, Dec 2005
2004
- Design Patterns for Reconfigurable ComputingIn IEEE Symposium on Field-Programmable Custom Computing Machines, Apr 2004
- Saliency on a chip: a digital approach with an FPGAThe Neuromorphic EngineerVolume 1, issue 2, Autumn 2004 , 2004