Engineering PapersSearch

SEARCH · Engineering Papers

Results for “codesign”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Neural Architecture Search is a powerful approach for automating model design, but existing methods struggle to accurately optimize for real hardware performance, often relying on proxy metrics such as bit operations. We present Surrogate Neural Architecture Codesign Package (SNAC-Pack), an integrated framework that automates the discovery and optimization of neural networks focusing on FPGA deployment. SNAC-Pack combines Neural Architecture Codesign's multi-stage search capabilities with the Resource Utilization and Latency Estimator, enabling multi-objective optimization across accuracy, FPGA resource utilization, and latency without requiring time-intensive synthesis for each candidate model. We demonstrate SNAC-Pack on a high energy physics jet classification task, achieving 63.84% accuracy with resource estimation. When synthesized on a Xilinx Virtex UltraScale+ VU13P FPGA, the SNAC-Pack model matches baseline accuracy while maintaining comparable resource utilization to models optimized using traditional BOPs metrics. This work demonstrates the potential of hardware-aware neural architecture search for resource-constrained deployments and provides an open-source framework for automating the design of efficient FPGA-accelerated models.

Weitz, Jason [UC, San Diego] (ORCID:00090004631535

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Neural architecture search (NAS) is a powerful approach for automating model design, but existing methods often optimize for accuracy alone or rely on proxy metrics such as bit operations (BOPs) that correlate poorly with hardware cost. This gap is particularly large for FPGA deployment, where cost is dominated by a multi-dimensional budget of lookup tables, DSPs, flip-flops, BRAM, and latency. We present the Surrogate Neural Architecture Codesign Package (SNAC-Pack), an open-source AutoML framework for hardware-aware neural architecture codesign and end-to-end FPGA deployment. SNAC-Pack runs a multi-objective global search with Optuna and NSGA-II, loading trials to a shared SQLite store that enables parallel workers across compute nodes. A hardware surrogate model outputs per-trial resource and latency estimates, avoiding the synthesis cost that would otherwise dominate the search loop. A local search stage then applies quantization-aware training (QAT) together with iterative magnitude pruning in a combined compression loop, after which the final model is synthesized to FPGA firmware via the hls4ml Python library. A YAML configuration and an optional agentic frontend let users run the pipeline on new datasets without modifying the framework. We demonstrate SNAC-Pack on jet classification at the Large Hadron Collider and superconducting qubit readout, discovering compact architectures that match or exceed strong baselines on the task metric while reducing FPGA resource utilization and, in the qubit readout case, reducing the design space exploration process from months of manual fine-tuning to hours of automated search.

Weitz, Jason [UC, San Diego]

Uncontrolled Learning: Codesign of Neuromorphic Hardware Topology for Neuromorphic Algorithms

Neuromorphic computing has the potential to revolutionize future technologies and our understanding of intelligence, yet it remains challenging to realize in practice. The learning-from-mistakes algorithm, inspired by the brain's simple learning rules of inhibition and pruning, is one of the few brain-like training methods. This algorithm is implemented in neuromorphic memristive hardware through a codesign process that evaluates essential hardware trade-offs. While the algorithm effectively trains small networks as binary classifiers and perceptrons, performance declines significantly with increasing network size unless the hardware is tailored to the algorithm. This work investigates the trade-offs between depth, controllability, and capacity—the number of learnable patterns—in neuromorphic hardware. This highlights the importance of topology and governing equations, providing theoretical tools to evaluate a device's computational capacity based on its measurements and circuit structure. The findings show that breaking neural network symmetry enhances both controllability and capacity. Additionally, by pruning the circuit, neuromorphic algorithms in all-memristive circuits can utilize stochastic resources to create local contrasts in network weights. Through combined experimental and simulation efforts, the parameters are identified that enable networks to exhibit emergent intelligence from simple rules, advancing the potential of neuromorphic computing.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Neural architecture codesign for fast physics applications

We develop a pipeline to streamline neural architecture codesign for physics applications to reduce the need for ML expertise when designing models for novel tasks. Our method employs neural architecture search and network compression in a two-stage approach to discover hardware efficient models. This approach consists of a global search stage that explores a wide range of architectures while considering hardware constraints, followed by a local search stage that fine-tunes and compresses the most promising candidates. We exceed performance on various tasks and show further speedup through model compression techniques such as quantization-aware-training and neural network pruning. We synthesize the optimal models to high level synthesis code for FPGA deployment with the hls4ml library. Additionally, our hierarchical search space provides greater flexibility in optimization, which can easily extend to other tasks and domains. We demonstrate this with two case studies: Bragg peak finding in materials science and jet classification in high energy physics, achieving models with improved accuracy, smaller latencies, or reduced resource utilization relative to the baseline models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Bridging the Gap Between LLMs and LNS with Dynamic Data Format and Architecture Codesign

Deep neural networks (DNNs) have achieved tremendous success in the past few years. However, their training and inference demand exceptional computational and memory resources. Quantization has been shown as an effective approach to mitigate the cost, with the mainstream data types reduced from FP32 to FP16/BF16 and recently FP8 in the latest NVIDIA H100 GPUs. With increasingly aggressive quantization, however, the conventional floating-point formats suffer from limited precision in representing numbers around zero. Recently, NVIDIA demonstrated the potential of using a Logarithmic Number System (LNS) for the next generation of tensor cores. While LNS mitigates the hurdles in representing small numbers, in this work we observed a mismatch between LNS and the emerging Large Language Models (LLM), where LLM exhibits significant outliers when directly adopting the LNS format. In this paper, we present a data-format/architecture codesign to bright this gap. On the format side, we propose a dynamic LNS format to flexibly represent outliers at a higher precision, by exploiting asymmetry in the LNS representation and identifying outliers through a per-vector basis. On the architecture side, for demonstration, we realize the dynamic LNS format in a systolic array, which can handle the irregularity of the outliers at runtime. We implement our approach on an Alveo U280 FPGA as a prototype. Experimental results show that our design can effectively handle the outliers and resolve the mismatch between LNS and LLM, contributing to an accuracy improvement of 15.4% and 16% over the floating-point and the original LNS baselines, using four state-of-the-art LLM models. Our observation and design lay a solid foundation for the large-scale adoption of the LNS format in the next-generation deep learning hardware.

Haghi, Pouya

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Neural Architecture Search is a powerful approach for automating model design, but existing methods struggle to accurately optimize for real hardware performance, often relying on proxy metrics such as bit operations. We present Surrogate Neural Architecture Codesign Package (SNAC-Pack), an integrated framework that automates the discovery and optimization of neural networks focusing on FPGA deployment.

Weitz, Jason [UC, San Diego]

Optimizing spin qubit coherence through materials codesign

The evolution of defect-based spin qubit systems is currently transitioning from fundamental studies and proof-of-concept demonstrations into applications in the burgeoning field of quantum technology. Within this context, new challenges emerge, in particular, the need to understand and engineer the fundamental materials that form the hardware building blocks critical for the scalability and wide-scale adoption of such technologies. While earlier discussions have often focused on qubits within idealized systems, major limitations on spin coherence and optical properties arise from effects imposed by the nonideality of the surrounding host matrix. Decoherence can stem from a variety of sources, including other qubits, nuclear spins, and parasitic point- and extended defects, which interact with the qubit via magnetic and electric fields, photons, phonons, and strain. In this article, we focus on the relevant sources and mechanisms through which decoherence occurs and provide potential mitigation strategies via the synergistic integration of first-principles simulations and materials synthesis and engineering. We aim to provide a tangible link between material properties and material functions thereby enabling materials-by-design.

ab initio simulations

End-to-end codesign of Hessian-aware quantized neural networks for FPGAs

Here, we develop an end-to-end workflow for the training and implementation of co-designed neural networks (NNs) for efficient field-programmable gate array (FPGA) hardware. Our approach leverages Hessian-aware quantization of NNs, the Quantized Open Neural Network Exchange intermediate representation, and the hls4ml tool flow for transpiling NNs into FPGA firmware. This makes efficient NN implementations in hardware accessible to nonexperts in a single open sourced workflow that can be deployed for real-time machine-learning applications in a wide range of scientific and industrial settings. We demonstrate the workflow in a particle physics application involving trigger decisions that must operate at the 40-MHz collision rate of the CERN Large Hadron Collider (LHC). Given the high collision rate, all data processing must be implemented on FPGA hardware within the strict area and latency requirements. Based on these constraints, we implement an optimized mixed-precision NN classifier for high-momentum particle jets in simulated LHC proton-proton collisions.

47 OTHER INSTRUMENTATION

Characterizing GPU Energy Usage in Exascale-Ready Portable Science Applications

We characterize the GPU energy usage of two widely adopted exascale-ready applications representing two classes of particle and mesh solvers: (i) QMCPACK, a quantum Monte Carlo package, and (ii) AMReX-Castro, an adaptive mesh astrophysical code. We analyze power, temperature, utilization, and energy traces from double-/single (mixed)-precision benchmarks on NVIDIA’s A100 and H100 and AMD’s MI250X GPUs using queries in NVML and rocm_smi_lib, respectively. We explore application-specific metrics to provide insights on energy vs. performance trade-offs. Our results suggest that mixed-precision energy savings range between 6–25% on QMCPACK and 45% on AMReX-Castro. Also, we found gaps in the AMD tooling used on Frontier GPUs that need to be understood, while query resolutions on NVML have little variability between 1 ms-1 s. Overall, application level knowledge is crucial to define energy-cost/science-benefit opportunities for the codesign of future supercomputer architectures in the post-Moore era.

Godoy, William [ORNL] (ORCID:0000000225905178)

Simultaneous optimal system and controller design for multibody systems with joint friction using direct sensitivities

Abstract Real-world multibody systems are often subject to phenomena like friction, joint clearances, and external events. These phenomena can significantly impact the optimal design of the system and its controller. This work addresses the gradient-based optimization methodology for multibody dynamic systems with joint friction using a direct sensitivity approach. The Brown–McPhee model has been used to characterize the joint friction in the system. This model is suitable for the study due to its accuracy for dynamic simulation and its compatibility with sensitivity analysis. This novel methodology supports codesign of the multibody system and its controller, which is especially relevant for applications like robotics and servo-mechanical systems, where the actuation and design are highly dependent on each other. Numerical results are obtained using a software package written in Julia with state-of-the-art libraries for automatic differentiation and differential equations. Three case studies are provided to demonstrate the attractive properties of simultaneous optimal design and control approach for certain applications.

Verulkar, Adwait

Recent progress on coarse graining simulations

We focus on coarse graining simulations based on the primary conservation equations, effectively codesigned physics and algorithms, and low-Mach-number corrected (LMC) hydrodynamics. Simulation methods involve LANL’s x-Radiation-Adaptive-Grid-Eulerian Large-Eddy Simulation, Besnard-Harlow-Rauenzahn (BHR) Reynolds-Averaged Navier-Stokes (RANS) approach, and Dynamic BHR – a paradigm bridging RANS and LES. A relevant question addressed relates to whether 3D RANS and RANS/LES hybrids – the industry standards for aerospace and automotive research, are presently relevant for practical variable-density applications involving shocked and accelerated interface instabilities. Furthermore, recent simulations of the GaTECH inclined mixing-layer shock-tube and NIF ICF-capsule experiments are used to demonstrate issues, challenges, and potential for 3D coarse grained LMC simulation strategies for robustly simulating complex transitional and coupled hydrodynamics-multiphysics with coarser resolution. Present LES readiness to provide accurate predictions at scale is demonstrated – whereas 3D RANS and RANS/LES bridging do not appear impactful in this context.

42 ENGINEERING

Amine Structure Governs Corrosion Rates of Copper Catalysts in Electrochemical Reactive Capture of CO 2

Reactive capture of CO 2 (RCC) offers an integrated approach that combines CO 2 capture with its direct electrochemical conversion, eliminating the need for CO 2 release from the capture agent. By avoiding the pH, pressure, and temperature swings required for the release step, RCC has the potential to reduce both energy consumption and capital costs compared to the conventional sequential process of CO 2 capture, release, concentration, and conversion. Amines, widely used in industrial CO 2 capture, face challenges in RCC systems due to their incompatibility with transition metal catalysts as well as their tendency to promote electrode corrosion and parasitic hydrogen evolution. Identifying suitable combinations of amines and catalysts is therefore critical to enabling integrated CO 2 capture and conversion. Here, this work systematically investigates the performance of four primary and four secondary amines for RCC on polycrystalline Cu catalysts. Among the eight tested amines, only dimethylamine showed no measurable Cu corrosion near the open circuit potential. In contrast, ammonia, methylamine, ethylamine, monoethanolamine, diethylamine, diethanolamine, and piperazine all induced Cu corrosion. Corrosion rates correlate with the pK a and steric hindrance of the amines, highlighting key parameters for catalyst–amine codesign. Grand canonical DFT calculations indicate a correlation between the adsorption strength of protonated amines, their pK a , and the extent of Cu corrosion, suggesting that both the surface binding of protonated amines and the lability of their protons play critical roles in corrosion acceleration near open circuit potentials. These finding suggest that amines with high pK a values and weak binding of their protonated forms to Cu surfaces are preferred, as they offer better corrosion resistance.

Choi, Jounghwan [Univ. of California, Los Angeles,

Engineering Assembly Kinetics and Line Roughness in Solvent Vapor-Annealed Block Copolymer/Homopolymer Blends

Block copolymer (BCP) directed self-assembly (DSA) is a promising route to enhance lithography resolution by multiplying nanopattern density and reducing feature roughness. Eliminating kinetically trapped self-assembly defects requires fast self-assembly. However, acceleration strategies like solvent vapor annealing or homopolymer blending broaden domain interfaces, implying a trade-off in increased feature roughness. In this work, we experimentally investigate this apparent dilemma between self-assembly kinetics and line roughness for solvent vapor-annealed thin films of a lamellar poly(styrene-block-2-vinylpyridine) (PS-b-P2VP) BCP blended with PS and P2VP homopolymers. Binary blends with PS or P2VP homopolymers and ternary blends incorporating both in equal weight fractions were solvent vapor annealed using acetone, a near-neutral solvent for PS and P2VP, followed by P2VP-selective vapor-phase infiltration with alumina (AlOx) and polymer etching. Binary blends with P2VP exhibit a modest kinetic enhancement but also higher line-edge and -width roughness due to the increased frequency of P2VP protrusions and bridge defects in the alumina line patterns. In contrast, binary blends with PS self-assemble noticeably faster, while domain asymmetry from the added PS homopolymer reduces roughness by curbing the number of alumina protrusions and bridge defects. Ternary blends maintain DSA line patterns across a wider composition window and, at higher homopolymer loadings, reduce roughness at length scales near the lamellar period, consistent with a reduced impact of intradomain compositional fluctuations. These findings provide important insights for codesigning blend compositions and process flows to achieve high-resolution, defect free patterns with minimal roughness through BCP DSA.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH