Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Co-design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Quantum Computing Strategy 2026

Quantum computing (QC) is a rapidly maturing technology with the potential for revolutionary impacts on stockpile stewardship science and national security. Recent developments in fault-tolerant architectures have compressed vendor roadmaps, and predictions of a production-ready quantum computer by the mid-2030s are becoming increasingly credible. This strategy provides a roadmap for integrating QC into the Advanced Simulation and Computing (ASC) program by investing in four strategic focus areas: 1. Develop Capabilities in Mission-Relevant Quantum Applications: ASC will prioritize developing quantum-ready applications in mission areas that have shown significant promise for quantum advantage, including simulations of materials in extreme environments, nuclear dynamics, solving linear and nonlinear partial differential equations, and uncertainty quantification. These applications directly support stockpile stewardship science and modernization objectives. 2. Conduct R&D in Algorithms, Software, and Hardware: Sustained research into quantum algorithms, robust software tools, and quantum hardware is essential. ASC will develop efficient quantum algorithms; invest in quantum compilers, debuggers, and performance tools; and explore specialized quantum hardware tailored to NNSA’s unique requirements. 3. Engage with Vendors and Partners: Early and active collaboration with commercial quantum hardware vendors and academic partners is critical. Through testbeds, co-design agreements, and quantum demonstration facilities, ASC will influence hardware design, gain early access to emerging technologies, and ensure that quantum platforms evolve to meet mission needs. 4. Build Knowledge, Experience, and Workforce: Expanding and upskilling the quantum-trained workforce is essential to long-term success. This includes hiring, internal training, university outreach, and postdoctoral support to ensure ASC maintains the expertise required to operate, program, and integrate quantum systems as they become available. While quantum computing will never replace classical computing, it has the potential to solve certain problems with speed and accuracy that would be unachievable using any conceivable classical high-performance computing (HPC) system. By investing strategically in QC, ASC will help propel the emergent QC industry, maintain U.S. technological leadership, ensure mission readiness, and position itself to rapidly adopt quantum technologies as they mature.

97 MATHEMATICS AND COMPUTING↗

Improved Charge Sensing on a SiMOS Double Quantum Dot using a Cryogenic Skipper Readout ASIC (Quandarum)

Major outstanding questions in high-energy physics such as the nature of dark matter and the existence of interactions beyond the standard model require new measurement techniques which are extremely sensitive to minute electromagnetic fields. An array of entangled spin qubits is a promising system for building novel detectors due to its combination of sensitivity and controllability. CMOS-based electron spin qubits, which have demonstrated the operational requirements for fault-tolerant quantum computing [1], offer a particular opportunity due to their compatibility with classical electronics, which allows the leveraging of decades of development of low-noise cryogenic detectors for physics. In this work, we combine a SiMOS double-quantum dot device architecture with a state-of-the-art cryoelectronic readout circuit [2-3] aimed to demonstrate improved charge readout using a single-electron transistor (SET). We identify the design characteristics for an SET that facilitate the use of on-chip classical electronics as a low-power, high-bandwidth first amplification stage and explore opportunities for sensor-readout co-design to minimize noise. This is the first of a series of steps to demonstrate high-fidelity readout of a large array of spin qubit with enough sensitivity to probe processes of interest for the investigation of beyond-standard-model physics.

Quinn, Adam [Fermilab]↗

28nm front end ASIC and 12” LGADs for 3D integration

The 3DIntSenS Collaboration—a joint effort between SLAC, Fermilab, and LLNL—is developing enabling technologies for next-generation radiation imaging detectors that combine ultra-fine spatial resolution (≈10 μm) with precision timing (<20 ps), while maintaining low power <1 W/cm2 and high data throughput. The approach leverages 3D integration between advanced CMOS readout ASICs and finely pixelated LGAD sensors to achieve the performance and scalability required for large-area, high-rate applications. High-granularity, precision-timing detectors are essential for scientific advances in HEP, NP, BES, and FES, but widespread adoption is limited by the cost and complexity of 3D integration. To close this gap, the collaboration is developing LGAD sensors compatible with 12-inch commercial CMOS processes, enabling cost-effective integration with high-performance ASICs under development. We present the design and results from a 28 nm CMOS ASIC prototype, including a low-jitter front end, and in-pixel TDC demonstrating sub-10 ps timing resolution. We also report on the co-design and characterization of reticle-scale LGAD sensors with 50 μm and 100 μm pixels and introduce the next 10k-pixel ASIC designed for full 3D integration. These advances represent a critical step toward scalable, high-resolution radiation imaging systems for future scientific instrumentation.

England, Troy [Fermilab] (ORCID:0000000154405255)↗

Systems-To-Atoms (S2A): enabling hydrogen for climate security

The project addresses a critical gap in hydrogen infrastructure by integrating system-level energy models with atomic-scale material simulations in a unified Systems-to-Atoms (S2A) framework. The motivation stems from the need to develop efficient, cost-effective, and durable hydrogen transport and utilization technologies to support decarbonization of hard-to-electrify sectors such as heavy-duty transportation. Current system models lack awareness of material performance mechanisms, while material-scale models do not account for system-level usage and variability. To bridge this divide, the team developed a co-simulation capability linking techno-economic analyses, reactor/process-flow modeling, and molecular-scale catalysis simulations. Applied to hydrogen delivery in California, the framework enabled comparative evaluations of compressed, cryogenic, and liquid organic hydrogen carrier (LOHC) pathways, highlighting how catalyst operation and unit process efficiency influence overall performance. The results demonstrate that no single material or transport mode is universally optimal; instead, heterogeneous solutions tuned to specific operational contexts deliver better performance. The project delivers a new capability for cross-scale material co-design, advancing hydrogen infrastructure readiness and informing DOE and LLNL missions in climate and energy resilience.

organic↗

Accelerating computing for the future electric grid (CRADA Final Report)

As a participant in the Cyclotron Road Lab-Embedded Entrepreneurship Program (LEEP), Vellex Computing, Inc. has successfully validated the "Vellex Computing Stack," a breakthrough Analog Neural Computer (ANC) specifically designed for high-performance edge optimization. This project achieved critical milestones in mixed-signal circuit stability and software-hardware co-design, directly addressing national priorities in semiconductor resiliency. The success of this work is deeply rooted in the support from the Cyclotron Road LEEP, which provided the essential "hard tech" runway—funding, mentorship, and access to Lawrence Berkeley National Laboratory’s world-class characterization facilities—allowing Vellex to overcome the "Valley of Death" often faced by deep-tech hardware startups. By leveraging LBNL’s advanced testing infrastructure, Vellex was able to rigorously benchmark the ANC architecture against state-of-the-art digital solutions, a feat that would have been resource-prohibitive independently. This collaboration has not only advanced American leadership in analog computing but has also matured Vellex’s technology to a stage ripe for private sector commercialization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Towards a Verifiable Domain-Specific Language for Hardware-Accelerated Stencils

Defining a domain-specific language (DSL) that supports vector-calculus abstractions eases the porting of partial differential equation (PDE) solvers to specialized architectures. Sufficiently high-level abstractions empower users to express universal laws with sufficient generality that the laws must always hold true within their domain of validity. A broad class of PDE solvers employs stencil-based algorithms, the target domain of Berkeley Lab's stencil accelerator chip co-design project. First released as open-source in January 2026, the Formal software framework lays a foundation for defining an embedded DSL based on composable operators that implement mimetic numerical methods -- stencil algorithms that guarantee satisfaction of discrete versions of important vector calculus theorems. The Formal DSL will be the frontend to a new class of stencil-PDE accelerators developed jointly by LBNL, UHCL, and UC Berkeley through the DOE Competitive Portfolios for Computer Science Project. This offers the potential of an order of magnitude acceleration for this important category of computational methods to serve the DOE mission. Future work on the Formal DSL will facilitate software verification via type-safe templates that enable problem-specific correctness proofs relying upon generic function theory and carefully crafted unit tests.

Rouson, Damian↗

Dependable classical-quantum computing systems engineering

Increasing evidence suggests quantum computing (QC) complements traditional High-Performance Computing (HPC) by leveraging its unique capabilities, leading to the emergence of a new, hybrid paradigm, QHPC. However, this integration introduces new challenges, with dependability–defined by reproducibility, resiliency, and security and privacy–emerging as a central concern for building trustworthy systems that provide an advantage to the users. This paper proposes a framework for dependable QHPC system design, organized around these three pillars. We identify integration challenges, anticipate roadblocks, and highlight productive synergies across QC, HPC, cloud platforms, and network security. Drawing from both classical computing principles and quantum-specific insights, we present a roadmap for co-design that supports robust hybrid architectures. Our approach offers concrete metrics for assessing dependability, provides design guidance for engineers working at the QC-HPC interface, and surfaces new engineering questions around complexity, scale, and fault tolerance. Ultimately, designing for dependability is key to realizing practical, scalable QHPC systems and accelerating the broader quantum ecosystem capable of translating quantum promises into actual application delivery.

HPC↗

Structural Characterization of Linker Shielding in ADC Site-Specific Conjugates

Background/Objectives: Antibody–Drug Conjugates (ADCs) have rapidly evolved from early, rudimentary conjugates to highly targeted and precisely engineered molecules. Despite notable clinical successes, ADCs continue to face significant challenges, including aggregation and high hydrophobicity driven by high drug-to-antibody ratios (DARs), premature payload release, dose-limiting toxicities, and suboptimal pharmacokinetics. While site-specific linker–payload conjugation has improved ADC homogeneity and stability, the structural basis of antibody–linker interactions at specific sites remains underexplored. Methods: In this work, we present the crystal structures of trastuzumab Fab and Fc domains site-specifically conjugated with a cleavable linker–payload. Results: Our findings suggest that pockets within both Fab and Fc regions may interact with and shield the linker portion of the conjugate. Conclusions: These insights highlight the previously underappreciated potential of structure-based design to drive the optimization of ADC linker chemistry and facilitate the co-design of bespoke linker–payloads tailored to individual antibody conjugation sites.

Jaime-Garza, Maru [Discovery Chemistry, Merck & Co↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

FOS: Computer and information sciences↗

Advanced-Research-on-Integrated-Energy-Systems-Based Analysis to Support Resilient System Upgrades: Energy to Communities Energyshed In-Depth Partnership with Molokai, Hawaii

The Molokai, Hawaii, Energy to Communities (E2C) Energyshed project represents a collaborative effort between the National Laboratory of the Rockies, Shake Energy Collaborative, the Molokai Clean Energy Hui, Sustainable Molokai, and Ho'ahu Energy Cooperative Molokai to advance Molokai's Community Energy Resilience Action Plan (CERAP). Supported by Hawaiian Electric Company and the Hawaii State Energy Office, the initiative aims to develop a community-defined portfolio of renewable energy solutions that enhance energy resilience while aligning with the Hawaiian Electric Integrated Grid Plan (IGP) and Molokai's energy goals. Phase 1 focused on technical analyses and community engagement to co-design feasible energy scenarios. Challenges such as grid upgrades, storage sizing, and inverter ride-through standards were addressed to align technical and operational requirements with community preferences. The project equips Molokai with actionable data and insights to implement energy initiatives while ensuring resilient and culturally informed solutions. Future efforts aim to finalize project designs, secure interconnection agreements, and deploy energy projects that reflect community priorities and technical feasibility.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Simulation of 100-300 GHz solid-state harmonic sources

Accurate and efficient simulations of the large-signal time-dependent characteristics of second-harmonic Transferred Electron Oscillators (TEO's) and Heterostructure Barrier Varactor (HBV) frequency triplers have been obtained. This is accomplished by using a novel and efficient harmonic-balance circuit analysis technique which facilitates the integration of physics-based hydrodynamic device simulators. The integrated hydrodynamic device/harmonic-balance circuit simulators allow TEO and HBV circuits to be co-designed from both a device and a circuit point of view. Comparisons have been made with published experimental data for both TEO's and HBV's. For TEO's, excellent correlation has been obtained at 140 GHz and 188 GHz in second-harmonic operation. Excellent correlation has also been obtained for HBV frequency triplers operating near 200 GHz. For HBV's, both a lumped quasi-static equivalent circuit model and the hydrodynamic device simulator have been linked to the harmonic-balance circuit simulator. This comparison illustrates the importance of representing active devices with physics-based numerical device models rather than analytical device models.

NONLINEAR CIRCUITS↗

Electronic Design Automation: Integrating the Design and Manufacturing Functions

As the complexity of electronic systems grows, the traditional design practice, a sequential process, is replaced by concurrent design methodologies. A major advantage of concurrent design is that the feedback from software and manufacturing engineers can be easily incorporated into the design. The implementation of concurrent engineering methodologies is greatly facilitated by employing the latest Electronic Design Automation (EDA) tools. These tools offer integrated simulation of the electrical, mechanical, and manufacturing functions and support virtual prototyping, rapid prototyping, and hardware-software co-design. This report presents recommendations for enhancing the electronic design and manufacturing capabilities and procedures at JSC based on a concurrent design methodology that employs EDA tools.

Bachnak, Rafic↗

Evaluation of the Telecommunications Protocol Processing Subsystem Using Reconfigurable Interoperable Gate Array

The current implementation of the Telecommunications Protocol Processing Subsystem Using Reconfigurable Interoperable Gate Arrays (TRIGA) is equipped with CFDP protocol and CCSDS Telemetry and Telecommand framing schemes to replace the CPU intensive software counterpart implementation for reliable deep space communication. We present the hardware/software co-design methodology used to accomplish high data rate throughput. The hardware CFDP protocol stack implementation is then compared against the two recent flight implementations. The results from our experiments show that TRIGA offers more than 3 orders of magnitude throughput improvement with less than one-tenth of the power consumption.

protocol processing hardware↗

Wind Energy Accomplishments and Year-End Performance Report: Fiscal Year 2024

As the largest source of clean, renewable power generation in the United States and one of the fastest growing sources of new electricity supply, wind energy will play a large role in the nation's energy future. In Fiscal Year (FY) 2024, scientists, engineers, analysts, and support professionals at the U.S. Department of Energy's (DOE's) National Renewable Energy Laboratory (NREL) worked to accelerate the pace of innovation in wind energy science and technology, advance grid systems integration, and develop sustainable solutions to deployment challenges. Much of NREL's research, development, and deployment work aligns with addressing the Grand Challenges of Wind Energy. Beginning in 2019, DOE's Wind Energy Technologies Office partnered with the International Energy Agency to identify the barriers to greater wind energy deployment and related research gaps. The world's leading wind energy scientists and engineers identified five research areas as critical to advancing wind energy deployment: wind atmospheric science, wind turbine systems, wind plants and grid, environmental co-design, and social science. In FY 2024, NREL's accomplishments helped narrow the research gaps in these critical areas. This report provides details on those accomplishments.

accomplishments↗

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

Schulte, Jan-Frederik [Purdue U.] (ORCID:000000034↗

On-chip probabilistic inference for charged-particle tracking at the sensor edge

Modern scientific instruments operate under increasingly extreme constraints on bandwidth, latency, and power. Inference at the sensor edge determines experimental data collection efficiency by deciding which information to save for further analysis. Particle tracking detectors at the Large Hadron Collider exemplify this challenge: pixelated silicon sensors generate rich spatiotemporal ionization patterns, yet most of this information is discarded due to data-rate limitations. Concurrently, advancements in co-design tools provide rapid turn-around for incorporating machine learning into application-specific integrated circuits, motivating designs for particle detectors with new integrated technologies. We demonstrate that neural networks embedded in the front-end electronics can infer charged-particle kinematic parameters from a single silicon layer. We regress hit positions and incident angles with calibrated uncertainties, while satisfying stringent constraints on numerical precision, latency, and silicon area. Our results establish a path toward probabilistic inference directly at the edge, opening new opportunities for intelligent sensing in high-rate scientific instruments.

Das, Arghya Ranjan [Purdue U.] (ORCID:000000018451↗

Efficient Routing of Quantum LDPC Codes on Programmable 2D Toric Architectures

Quantum low-density parity-check codes are promising candidates towards scalable fault-tolerant quantum computation. Among these, bivariate bicycle (BB) codes offer superior encoding rates and large code distance compared to surface codes. However, their requirement on long-range stabilizer measurements poses significant challenges for implementation on realistic hardware with limited connectivity, such as superconducting circuit platforms. In this work, we introduce a novel hardware-software co-design that leverages a programmable communication network architecture to address these limitations. Our approach utilizes a 2D toric network of oscillators as a flexible communication fabric linking qubits at each site. Such architecture significantly reduces the number of long-range couplers required from O ( n ) to O (√ n ). Dual-rail qubits, along with native gates including Swap-Wait-Swap gates and beamsplitter SWAPs, ensure that long-range two-qubit gates can be executed with high fidelity and low latency. To further enhance performance, our qubit layout and routing algorithm utilize symmetries of the codes and enable maximum parallelism for long-range two-qubit gates, maintaining a low syndrome extraction cycle duration and scalability over the code length. We perform circuit-level simulation with realistic noise modeling based on experimental hardware parameters, observing an logical error rate per logical qubit per cycle of 3.06% for [[18,4,4]] BB code, 2.6× less than the existing experimental result. These findings provide a practical roadmap and identify key technological advancements needed to achieve low-overhead fault-tolerant quantum computing at scale.

Liu, Kun [Yale Univ., New Haven, CT (United States↗