Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “control co-design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Co-Design of Marine Energy Converters for Autonomous Underwater Vehicle Docking and Recharging - Software and Data

Software and testing data from the OH Hinsdale Wave lab for DOE-funded project on Co-Design of Marine Energy Converters for Autonomous Underwater Vehicle Docking and Recharging. This project will perform foundational research and testing to accelerate the sector-wide development and deployment of marine energy converters to provide Power-At-Sea. Specifically, we seek to overcome known challenges and knowledge gaps for the successful co-design of coupled Wave Energy Converter (WEC)-Autonomous Underwater Vehicles (AUV) systems; systems designed and tested for WEC array system health and environmental monitoring applications. This project brings together an experienced, multi-institution, and multi-disciplinary team to focus on the co-design of marine energy (ME) technologies and AUV docking systems, including multi-body hydrodynamic modeling, active control, autonomy, and hardware interfaces necessary to enable new WEC-focused understanding, and allow for robust and ubiquitous AUV docking and recharging in real-world conditions.

16 TIDAL AND WAVE POWER↗

Report on use of Inoculants in Missile Application Alloys

This report documents the status of current inoculant research relevant to missile application alloys and MTCR control language. The information is intended to provide data on current inoculants for us determining the current state of development and identifying potential research directions. Although there has been significant scientific research into the development and synthesis of inoculants, their current availability is limited is traditional powder inoculants employed during casting processes. However, research continues the development of complex oxides, ribbon materials, high entropy alloys, and other inoculant product forms, including the use of inoculants in the melt pools produced during additive manufacturing. Research to date has focused primarily on aluminum and steel alloys with emphasis on refining grain structures and evolving equiaxed morphologies while increasing strength and castability. The primary inoculants in steel and cast irons include TiN, SiC, FeSi75, and Ce which have increased strength properties. Chief inoculants for Al alloys often include TiC, SiC, Al3Sc(x) and TiB2 to aid in precipitation and refinement. Ti and Ni alloys have fewer research activities involving inoculants, although TiN, TiB, ZrN and LaB6 (for Ti alloys) and WC, Co3FeNb2, and CrFeNb (for Ni alloys) have been used. Sic, Al2O3, Mg and Ti are key inoculants for Mg alloys. Multiple cast alloys from each of the material classes demonstrated increased strength and performance properties using inoculants, with several approaching requirements applicable to missile service environments. The continued evolution of advanced manufacturing capabilities is making it easier to produce high temperature near net shape structural materials using inoculant powders. These shapes may include the geometric shapes addressed within the MTCR (tubes and limited wall thicknesses). The use of inoculants may enable further development of high temperature alloys into near net shapes traditionally produced via casting processes due to limited ductility. This may decrease material and manufacturing costs. In addition, inoculation provides controlled kinetics and achievable chemical segregation that enables potential for far-from equilibrium thermodynamic microstructures and chemistries that could provide new metastable alloy states and subsequent properties to address co-design engineering constraints, including needs for increased strength and ductility. It is recommended that specific material combinations within these alloy classes be carefully watched as the materials evolve, with controls aimed at those having material properties above current MTCR levels. This specifically includes the use of refractory inoculants in alloys, and the application of inoculants in high strength and high temperature alloys via additive manufacturing processes, with care to link capabilities to product forms similar to the current requirements on tube geometries and material feed stocks. The continued development of nanoparticle inoculants will increase strength and ductility of high strength castings and additive manufactured metallic components. For example, adding inoculants into the casting of maraging steels and other precipitation strengthened alloys may drastically elevate mechanical properties above the control limit of current regulations.

36 MATERIALS SCIENCE↗

MASK4 Test Campaign for Sandia WaveBot Device

This data and report details the findings from a wave tank test focused on production of useful work of a wave energy converter (WEC) device. The experimental system and test were specifically designed to validate models for power transmission throughout the WEC system. Additionally, the validity of co-design informed changes to the power take-off (PTO) were assessed and shown to provide the expected improvements in system performance. These data describe the "MASK4" wave tank test of the Sandia WaveBot device. The WaveBot device has been tested a number of times in different permutations at the US Navy's Maneuvering and Sea Keeping (MASK) basin. Each test in this series is referred to as MASK1, MASK2, etc. The WaveBot device was first tested in one degree of freedom (heave) in 2016. This MASK1 test focused primarily on system identification and modeling. After MASK1, major modifications were performed to improve the overall real-time control and measurement system, improve the heave drive train, and add surge and pitch degrees of freedom. The second set of testing, which was broken up in to two stages: MASK2A and MASK2B, focused on bench testing and closed-loop control performance as well as nonlinear modeling. MASK3 then focused on multi-input, multi-output modeling and control for maximization of electrical power. The attached report presents the results from MASK4, which focuses on detailed modeling of the power conversion chain and validation co-design principles by way of the introduction of a magnetic spring. The test log, report, and data from the MASK4 test of the WaveBot augmented with a tunable magnetic spring. Processing codes can be found at the Github link below.

16 TIDAL AND WAVE POWER↗

Federated Learning and Differential Privacy: What might AI-Enhanced co-design of microelectronics learn?

Data is a valuable commodity, and it is often dispersed over multiple entities. Sharing data or models created from the data is not simple due to concerns regarding security, privacy, ownership, and model inversion. This limitation in sharing can hinder model training and development. Federated learning can enable data or model sharing across multiple entities that control local data without having to share or exchange the data themselves. Differential privacy is a conceptual framework that brings strong mathematical guarantee for privacy protection and helps provide a quantifiable privacy guarantee to any data or models shared. The concepts of federated learning and differential privacy are introduced along with possible connections. Lastly, some open discussion topics on how federated learning and differential privacy can tied to AI-Enhanced co-design of microelectronics are highlighted.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Cyber-Physical System: Design for Sustainability and Resilience

When considering the design tools needed in the transition from numeric models to pilot plant, cyber-physical systems (CPS) come to the forefront as a method to model complex integrated energy systems. CPS approach has proven to be valuable to identify opportunities for economically viable early adoption of integrated energy technologies. This tutorial will introduce the concepts and the roles of CPS in co-design to minimize risks for pilot plant and technology deployment. This tutorial will also layout basic requirements for the CPS development, which requires a highly interdisciplinary effort with expertise in sensors, hardware testing, real-time modeling, controls, and system integration.

Harun, Nor Farida↗

New trends in photonic switching and optical networking architectures for data centers and computing systems [Invited]

The rapid increases in data traffic coupled with user preferences are driving the data center and computing system service providers to offer energy-efficient, intelligent, flexible, cost-effective, high-capacity, and low-latency data services without added complexity to the users. Disaggregated heterogeneous reconfigurable computing systems realized by photonic switching and interconnects can enhance throughput and energy efficiency for artificial intelligence/machine learning (AI/ML) workloads, especially when aided by the AI/ML-enhanced control plane. Photonic switching and new optical networking architectures are expected to solve many of these challenging problems. This paper discusses new trends in photonic switching and optical network architectures for future data centers and computing systems summarized as follows: (1) flat reconfigurable disaggregated computing enabled by high-radix photonic switching and interconnects in data centers; (2) chiplet-based computing architectures empowered by embedded photonics toward heterogeneous reconfigurable computing; (3) nanosecond-scale photonic switching in data centers and computing systems; (4) AI/ML in self-driving, application-aware, and situation-aware data centers; (5) the emergence of flexible networking for cloud computing, edge computing, and split computing, as well as flexible networking for 5G/6G RF-optical networks; and (6) the deployment of embedded co-designed silicon photonics being considered for future data centers.

Yoo, S. J. Ben (ORCID:0000000274201871)↗

NeuroCoreX: An Open-Source FPGA-Based Spiking Neural Network Emulator with On-Chip Learning

Spiking Neural Networks (SNNs) are computational models inspired by the event-driven communication and connectivity patterns of biological neural circuits. They enable high energy efficiency and natural support for diverse architectures ranging from layered networks to small-world and graphstructured topologies. In this work, we introduce NeuroCoreX, an open-source, FPGA-based spiking neural network emulator that provides real-time, on-chip learning and flexible network organization. NeuroCoreX supports both feedforward sensory inputs streamed directly from sensors or PCs via UART and recurrent on-chip connectivity, enabling simultaneous processing and learning from external stimuli and internal network dynamics-capabilities rarely available in existing FPGA SNN platforms. The system implements a Leaky Integrate-and-Fire (LIF) neuron model with current-based synapses and supports pair-based STDP learning on both feedforward and recurrent synapses. A lightweight Python interface enables interactive configuration, live monitoring, weight read-back, and experiment control. Importantly, NeuroCoreX is tightly integrated with the SuperNeuroMAT simulator, allowing SNN models to be transferred seamlessly from software to hardware for hardware-in-the-loop development. By combining real-time plasticity, flexible connectivity, and an open-source VHDL implementation, NeuroCoreX provides an extensible and accessible platform for neuromorphic research, algorithm-hardware co-design, and energy-efficient edge intelligence.

Gautam, Ashish [ORNL]↗

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗

InterQnet: A Heterogeneous Full-Stack Approach to Co-Designing Scalable Quantum Networks

Quantum communications have progressed significantly, moving from a theoretical concept to small-scale experiments to recent metropolitan-scale demonstrations. As the technology matures, it is expected to revolutionize quantum computing in much the same way that classical networks revolutionized classical computing. Quantum communications will also enable breakthroughs in quantum sensing, metrology, and other areas. However, scalability has emerged as a major challenge, particularly in terms of the number and heterogeneity of nodes, the distances between nodes, the diversity of applications, and the scale of user demand. This article describes InterQnet, a multidisciplinary project that advances scalable quantum communications through a comprehensive approach that improves devices, error handling, and network architecture. InterQnet has a two-pronged strategy to address scalability challenges: InterQnet-Achieve focuses on practical realizations of heterogeneous quantum networks by building and then integrating first-generation quantum repeaters with error mitigation schemes and centralized automated network control systems. The resulting system will enable quantum communications between two heterogeneous quantum platforms through a third type of platform operating as a repeater node. InterQnet-Scale focuses on a systems study of architectural choices for scalable quantum networks by developing forward-looking models of quantum network devices, advanced error correction schemes, and entanglement protocols. Here, we report our current progress toward achieving our scalability goals.

Chung, Joaquin [Argonne] (ORCID:0000000173833810)↗

Co-design optimization of combined heat and power-based microgrids

With the emergent need for clean and reliable energy resources, hybrid energy systems, such as the microgrid, are widely adopted in the United States. A microgrid can consist of various distributed energy resources, for instance, combined heat and power (CHP) systems. Here, the CHP module is a distributed cogeneration technology that produces electricity and recaptures heat generated as a by-product. It is an energy-efficient technology converting heat that would otherwise be wasted to valuable thermal energy. For an optimal system configuration, this study develops a novel co-design optimization framework for CHP-based cogeneration microgrids. The framework provides the stakeholder with a method to optimize investments and attain resilient operations. The proposed co-design framework has a mixed integer programming (MIP) model that outputs decisions for both plant designs and operating controls. The microgrid considered in this study contains six components: the CHP, boiler, heat recovery unit, thermal storage system, power storage system, and photovoltaic plant. After solving the MIP model, the optimal design parameters of each component can be found to minimize the total installation cost of all components in the microgrid. Furthermore, the online costs from energy production, operation, maintenance, machine startup, and disruption-induced unsatisfied loads are minimized by solving the optimal control decisions for operations. Case studies based on designing a CHP-based microgrid with empirical data are conducted. Moreover, we consider both nominal and disruptive operational scenarios to validate the performance of the proposed co-design framework in terms of a cost-effective, resilient system.

42 ENGINEERING↗

SAN-Based Block Polymers as a Platform for Manufacturing Strong Isoporous Membranes

Ultrafiltration (UF) membranes are ubiquitous in water purification and bioprocessing. However, co-designing their mechanical and transport properties remains challenging because of the broad pore size distributions at the surface and within the bulk that result from nonsolvent-induced phase separation (NIPS) – their typical manufacturing process. These distributions influence the hydrodynamic resistance to water flow and the stress concentrations around the pores. Developing advanced UF membranes requires innovative molecular designs that offer control over the surface and bulk pores, as well as the mechanical properties of the load-bearing, polymer. Here, we introduce a platform for designing UF membranes by leveraging solution self-assembly of block polymers and chain architectures with pendant polar groups. The block polymers consist of a poly(styrene-co-acrylonitrile) hydrophobic block, which is known for its strength, and a poly(4-vinyl pyridine) hydrophilic block, which drives solution self-assembly. We focus on a series of block polymers with constant molecular weight, M n ≈ 115 kDa, SAN fraction, 75 wt.%, and varying acrylonitrile content, 0 to 40 mol%, to demonstrate that: (i) RAFT dispersion copolymerization of acrylonitrile and styrene provides a facile route to synthesize strong block polymers, (ii) incorporation of acrylonitrile into the hydrophobic block enhances membrane strength by facilitating chain entanglements and dipole-dipole interactions, and (iii) acrylonitrile alters the balance between membrane permeance and rejection, even when the membranes feature similar surface and bulk pores. Overall, our results provide insights into the molecular design of UF membranes with enhanced mechanical and separation properties, contributing to the development of materials for water and energy technologies.

deformation↗

Financial-technical co-design for capital-intensive, resource-responsive energy systems

Because of their capital-intensive operation, wind energy systems that are competitive in terms of the cost of the energy that they produce lead to risk-reward trade-offs that make their business cases less favorable than those of conventional energy generation technologies. However, wind energy systems tend to be designed to maximize energy production or minimize cost of energy rather than to maximize their business cases. In this work, we attempt to exploit designs specifically tailored to business cases. We develop a novel framework for analyzing energy systems that ties their design variables to monthly operating incomes using simple models and historical hourly market and resource data. Using this approach, we demonstrate that for a wind site with abundant wind resource in the California Independent System Operator market, we can control the trade-off between mean and 5th percentile monthly returns by choosing the specific power of the turbine at a fixed modeled initial capital cost. Our framework gives a measure of the risk-reward spectrum of energy generation assets that could be built at a given site with respect to the sub-annual resource/market variation.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Will Stochastic Devices Play Nice With Others in Neuromorphic Hardware?: There’s More to a Probabilistic System Than Noisy Devices

Achieving brain-like efficiency in computing requires a co-design between the development of neural algorithms, brain-inspired circuit design, and careful consideration of how to use emerging devices. The recognition that leveraging device-level noise as a source of controlled stochasticity represents an exciting prospect of achieving brain-like capabilities in probabilistic neural algorithms, but the reality of integrating stochastic devices with deterministic devices in an already-challenging neuromorphic circuit design process is formidable. Here, we explore how the brain combines different signaling modalities into its neural circuits as well as consider the implications of more tightly integrated stochastic, analog, and digital circuits. Further, by acknowledging that a fully CMOS implementation is the appropriate baseline, we conclude that if mixing modalities is going to be successful for neuromorphic computing, it will be critical that device choices consider strengths and limitations at the overall circuit level.

42 ENGINEERING↗

The Case for Co-Designing Model Architectures with Hardware

While GPUs are responsible for training the vast majority of state-of-the-art deep learning models, the implications of their architecture are often overlooked when designing new deep learning (DL) models. As a consequence, modifying a DL model to be more amenable to the target hardware can significantly improve the runtime performance of DL training and inference. In this paper, we provide a set of guidelines for users to maximize the runtime performance of their transformer models. These guidelines have been created by carefully considering the impact of various model hyperparameters controlling model shape on the efficiency of the underlying computation kernels executed on the GPU. We find the throughput of models with “efficient” model shapes is up to 39% higher while preserving accuracy compared to models with a similar number of parameters but with unoptimized shapes.

Yin, Junqi↗

QUCODE: End-to-End Qubit Co-Design

The design of a quantum computer can be broken down into different steps, e.g., the material science aspect of designing qubits and devices, considerations of controlling the state of the qubits and their environment, the computer science aspects of mapping algorithms to the available primitives of the quantum computer, and the programming of an application in terms of the available algorithms. Research in these areas is currently fairly isolated, and there is framework for an end-to-end design approach where a desired application informs the choice of materials for the qubits and their environment, and vice versa.We identify knowledge gaps and opportunities for research that builds on existing PNNL capabilities.

36 MATERIALS SCIENCE↗

Improved Charge Sensing on a SiMOS Double Quantum Dot using a Cryogenic Skipper Readout ASIC (Quandarum)

Major outstanding questions in high-energy physics such as the nature of dark matter and the existence of interactions beyond the standard model require new measurement techniques which are extremely sensitive to minute electromagnetic fields. An array of entangled spin qubits is a promising system for building novel detectors due to its combination of sensitivity and controllability. CMOS-based electron spin qubits, which have demonstrated the operational requirements for fault-tolerant quantum computing [1], offer a particular opportunity due to their compatibility with classical electronics, which allows the leveraging of decades of development of low-noise cryogenic detectors for physics. In this work, we combine a SiMOS double-quantum dot device architecture with a state-of-the-art cryoelectronic readout circuit [2-3] aimed to demonstrate improved charge readout using a single-electron transistor (SET). We identify the design characteristics for an SET that facilitate the use of on-chip classical electronics as a low-power, high-bandwidth first amplification stage and explore opportunities for sensor-readout co-design to minimize noise. This is the first of a series of steps to demonstrate high-fidelity readout of a large array of spin qubit with enough sensitivity to probe processes of interest for the investigation of beyond-standard-model physics.

Quinn, Adam [Fermilab]↗