Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Software frameworks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Analysis of Building Model Forecasts using Autonomous HVAC Optimization System for Residential Neighborhood

Heating, ventilation, and air conditioning (HVAC) systems account for the highest share of home energy consumption in the United States. Optimized HVAC control can provide thermal improved comfort to the occupants, improve energy efficiency, reduce energy cost, and support grid services. In this paper, we discuss a multi-agent and cloud-based software framework that has been deployed in occupied residential neighborhood. This system enables automatic data collection, learning, optimization, and dispatches signals to neighborhood devices. HVAC optimization is based on model predictive control (MPC). Since the operational performance of MPC depends on model forecasting accuracy, it is crucial to evaluate the model continuously and modify or retrain it as necessary. In this research, we developed an automated workflow to evaluate the performance of temperature and power forecasts based on measured data in the real world. This will provide researchers with a deeper understanding of the model and how it can be improved.

Lebakula, Viswadeep↗

ROAM: A Remotely Operated Accelerator Monitor

Monitoring accelerators in operation is a well known challenge due to the radiation environment. However, there are significant benefits in being able to deploy particular sensors in specific locations of accelerator enclosures for monitoring or troubleshooting purposes. Learning from experience at other labs, we used Commercial Off The Shelf (COTS) components and an open source robot control software framework (ROS) to build a remotely controlled robot platform including a standard suite of instruments such as cameras, LIDAR, and ultrasound, with the ability to incorporate other ad-hoc sensors for specific measurements, such as a Gamma radiation monitor. Special attention was given to the robot's ability to maintain safety in the high risk environment of an accelerator, with multiple failure contingencies in place to ensure collision free operation. Testing was performed to prove the platform's viability and showcase its capability for accelerator monitoring.

Thayer, Thomas C.↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

Bridging Python to Silicon: The SODA Toolchain

Systems performing scientific computing, data analysis, and machine learning tasks have a growing demand for application-specific accelerators that can provide high computational performance while meeting strict size and power requirements. However, the algorithms and applications that need to be accelerated are evolving at a rate that is incompatible with manual design processes based on hardware description languages. Agile hardware design tools based on compiler techniques can help by quickly producing an application-specific integrated circuit (ASIC) accelerator starting from a high-level algorithmic description. Here, we present the software-defined accelerator (SODA) synthesizer, a modular and open-source hardware compiler that provides automated end-to-end synthesis from high-level software frameworks to ASIC implementation, relying on multilevel representations to progressively lower and optimize the input code. Our approach does not require the application developer to write any register-transfer level code, and it is able to reach up to 364 giga floating point operations per second (GFLOPS)/W efficiency (32-bit precision) on typical convolutional neural network operators.

97 MATHEMATICS AND COMPUTING↗

Extending XACC for Quantum Optimal Control

Quantum computing vendors are beginning to open up application programming interfaces for direct pulse-level quantum control. With this, programmers can begin to describe quantum kernels of execution via sequences of arbitrary pulse shapes. This opens new avenues of research and development with regards to smart quantum compilation routines that enable direct translation of higher-level digital assembly representations to these native pulse instructions. In this work, we present an extension to the XACC system-level quantum-classical software framework that directly enables this compilation lowering phase via user-specified quantum optimal control techniques. This extension enables the translation of digital quantum circuit representations to equivalent pulse sequences that are optimal with respect to the backend system dynamics. Our work is modular and extensible, enabling third party optimal control techniques and strategies in both C++ and Python. We demonstrate this extension with familiar gradient-based methods like gradient ascent pulse engineering (GRAPE), gradient optimization of analytic controls (GOAT), and Krotov's method. Our work serves as a foundational component of future quantum-classical compiler designs that lower high-level programmatic representations to low-level machine instructions.

Nguyen, Thien↗

Pulsar Based Timing for Grid Synchronization

Existing synchronization systems in the power grid, such as the global positioning system, are susceptible to temporary or permanent failures due to various unpredictable and uncontrollable factors such as cyber-attack and electromagnetic interferences, thus affecting the accuracy and reliability of generated timing signal. In this article, a pulsar astronomy-based timing system is proposed to provide an alternative synchronization signal. Further, this clock will offer significant security improvements to power grid applications, such as a wide-area monitoring system, which depends on a precise timing signal. The hardware and software frameworks are described in detail. First, a high-speed sampling hardware platform is designed to collect signals from radio telescopes. Then a periodic pulse extraction method with three steps is proposed to process the pulsar signal, including polyphase filterbanks, incoherent de-dispersion, and sliding window folding. Lastly, three experiments are conducted to verify the effectiveness of the frameworks. The generated pulsar timing pulse is presented, and the factors affecting its accuracy are also discussed. The analysis results demonstrate that the pulsar signals can provide high-accurate timing pulses for grid synchronization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reliable edge machine learning hardware for scientific applications

Extreme data rate scientific experiments create massive amounts of data that require efficient ML edge processing. This leads to unique validation challenges for VLSI implementations of ML algorithms: enabling bit-accurate functional simulations for performance validation in experimental software frameworks, verifying those ML models are robust under extreme quantization and pruning, and enabling ultra-fine-grained model inspection for efficient fault tolerance. We discuss approaches to developing and validating reliable algorithms at the scientific edge under such strict latency, resource, power, and area requirements in extreme experimental environments. We study metrics for developing robust algorithms, present preliminary results and mitigation strategies, and conclude with an outlook of these and future directions of research towards the longer-term goal of developing autonomous scientific experimentation methods for accelerated scientific discovery.

Baldi, Tommaso↗

Smart Adaptive Structures for an Ocean Wave Energy Converter

Ocean wave energy converters face significant challenges including cost-effectiveness, minimizing maintenance requirements, and withstanding extreme conditions. However, by utilizing smart materials, these converters could overcome these challenges. Such energy harvesters could use dielectric elastomer generators to convert ocean wave energy into electricity through their dynamic straining. Conversely, by applying electricity to these generators, they become actuators - dielectric elastomer actuators - thereby enabling them to alter their stiffness and adapt to the ever-changing ocean energy environments. Such active adaptation could enhance the converter's ability to: reach resonance with ocean waves and protect itself from dangerous waves. This study utilizes numerical analyses through the COMSOL software framework to evaluate the potential energy that could be harvested by a conceptual ocean wave energy converter based upon dielectric elastomer generators/actuators. The converter is composed of an external hull (that is a hollow cylinder), an inertial mass (that is a hollow cylinder and concentric with the hull), and 'spokes' - made of dielectric elastomer generators/actuators - that connect the hull to the inertial mass. Results of the numerical analyses include those outcomes arising from the conceptual converter being simulated via a sinusoidal motion analogous to ocean waves. That motion, therefore, causes relative motion between the converter's hull and inertial mass thereby dynamically stretching the corresponding connecting elastomers. The stretching of the elastomers enables them to 'gain elastic strain energy' and is, therefore, considered to be the theoretical limit of possible electrical energy conversion for the dielectric elastomer generator/actuator spokes. Additionally, the elastomer material properties of the spokes were altered to simulate the actuation of those same elastomers; with overall strain energy being subsequently investigated. Ultimately this is a preliminary study exploring the ability of such smart materials - electricity-generating and self-actuating elastomers - to actively adapt an ocean wave energy converter's structure to address and overcome the aforementioned challenges.

active materials↗

Full event interpretation with machine-learning-based particle-flow reconstruction in the CMS detector

The particle-flow (PF) algorithm constructs a global description of each particle collision by producing a comprehensive list of final-state particles, and is central to event reconstruction in the CMS experiment at the CERN LHC. The existing PF implementation relies on physics-motivated heuristics and assumptions that can be replaced by machine-learning (ML) models trained directly on simulated data and naturally suited to modern graphics processing units (GPUs). A state-of-the-art ML-based PF (MLPF) reconstruction algorithm, implemented within the CMS software framework, is presented. The MLPF algorithm performs a learnable full-event reconstruction on GPUs, generalizes across detector conditions and collision energies, and replaces multiple modular reconstruction steps with a single unified model. Physics performance comparable to standard PF reconstruction is achieved in both simulation and data, with improved jet energy resolution and inference time. In simulated top quark-antiquark events under LHC Run-3 (2023-2024) conditions, the jet energy resolution improves by 10-20% for jets with transverse momentum between 30-100 GeV. Inference time is evaluated using simulated multijet events, with a median of $20\,\hbox {ms}$ per event on an Nvidia L4 GPU, compared to approximately $110\,\hbox {ms}$ for the standard CMS PF reconstruction.

Hayrapetyan, Aram [Yerevan Phys. Inst.]↗

mesoflow [SWR-22-56]

Mesoflow is a continuum scale simulation tool developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex surface morphologies (of catalysts/biomass particles among others) directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPU) compared to a single processor for representative problem sizes (2 million cell mesh).

Sitaraman, Hariswaran↗

Fugu v.0.1

SAND2021-15052 O Fugu provides a common software framework for designing and prototyping algorithms for spiking neuromorphic hardware and compiling to multiple hardware platforms. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Aimone, James↗

Multiplayer Engineering

INL engineers developed a special purpose software framework designed to establish real-time connections between multiple running Unity instances. The framework is developed in Unity, a game engine development platform. This framework programmatically attaches functions to engineering models, enabling users to interact with, observe, and make changes to virtual models. These interactions are broadcasted through websocket connections, and received by other engineers in real-time.

Woodruff, Nathan↗

pnnl/eqc

EQC is a software framework for smart distributed quantum computer management and organization for variational quantum algorithms. EQC utilizes smart real time management of multiple independent quantum processors to maximize variational quantum algorithm throughput, for improved converged performance and improved rate of convergance.

Central, PNNL Developer↗

AMReX v2024

The software framework, AMReX, supports the development of block-structured adaptive mesh refinement (AMR) algorithms for solving systems of partial differential equations. AMR reduces the computational cost and memory footprint compared to a uniform mesh while preserving the essential local descriptions of different physical processes in complex multiphysics algorithms. AMR uses a hierarchical representation of the solution at multiple levels of resolution where the solution on each level is defined on the union of data containers at that resolution. These data containers, which represent the solution over a logically rectangular subregion of the domain, can contain field data defined on a mesh, Lagrangian particles or combinations of both. In addition to these basic data types, AMReX supports a multilevel embedded boundary representation of complex geometry; linear solvers for cell-centered and nodal data; asynchronous I/O in a native format readable by ParaView, VisIt and yt; and interfaces to hypre and PETSc solvers. AMReX enables applications to run on distributed memory architectures with multicore CPUs and with GPU accelerators. AMReX uses a lightweight abstraction layer that effectively hides the details of the architecture from the application. The framework currently supports CUDA, HIP and SYCL for GPU acceleration and OpenMP for multi-core CPU architectures.

Almgren, Ann↗

Exascale-Enabled Models and Algorithms for Microelectronics Applications (MicroEleX) v1

The MicroEleX code package contains a variety of models and algorithms for physical modeling of microelectronic circuitry, including electrostatics, electrodynamics, superconducting physics, micromagnetics, multi-ferroic systems, and quantum transport. MicroEleX leverages the AMReX software framework to provide scalability on GPU-based supercomputing architectures. The code is open source and designed to be algorithmically flexible so developers can incorporate enhanced or customized physics.

Nonaka, Andy↗

pnnl/LAP

A software framework to study, from the performance and energy perspective, the efficacy of GPU-resident parallel Conjugate Gradient (CG) linear solver with different preconditioner options, including Gauss-Seidel, Jacobi, and incomplete Cholesky. We also propose a novel GPU-based preconditioner, in which the triangular solves are approximated by an iterative process

Swirydowicz, Kasia↗