Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗

ExaTN: Scalable GPU-Accelerated High-Performance Processing of General Tensor Networks at Exascale

We present ExaTN (Exascale Tensor Networks), a scalable GPU-accelerated C++ library which can express and process tensor networks on shared- as well as distributed-memory high-performance computing platforms, including those equipped with GPU accelerators. Specifically, ExaTN provides the ability to build, transform, and numerically evaluate tensor networks with arbitrary graph structures and complexity. It also provides algorithmic primitives for the optimization of tensor factors inside a given tensor network in order to find an extremum of a chosen tensor network functional, which is one of the key numerical procedures in quantum many-body theory and quantum-inspired machine learning. Numerical primitives exposed by ExaTN provide the foundation for composing rather complex tensor network algorithms. We enumerate multiple application domains which can benefit from the capabilities of our library, including condensed matter physics, quantum chemistry, quantum circuit simulations, as well as quantum and classical machine learning, for some of which we provide preliminary demonstrations and performance benchmarks just to emphasize a broad utility of our library.

97 MATHEMATICS AND COMPUTING↗

Benchmarking Optimizers for Qumode State Preparation with Variational Quantum Algorithms

Quantum state preparation involves preparing a target state from an initial system, a process integral to applications such as quantum machine learning and solving systems of linear equations. Recently, there has been a growing interest in qumodes due to advancements in the field and their potential applications. However there is a notable gap in the literature specifically addressing this area. This paper aims to bridge this gap by providing performance benchmarks of various optimizers used in state preparation with Variational Quantum Algorithms. We conducted extensive testing across multiple scenarios, including different target states, both ideal and sampling simulations, and varying numbers of basis gate layers. Our evaluations offer insights into the complexity of learning each type of target state and demonstrate that some optimizers perform better than others in this context. Notably, the Powell optimizer was found to be exceptionally robust against sampling errors, making it a preferred choice in scenarios prone to such inaccuracies. Additionally, the Simultaneous Perturbation Stochastic Approximation optimizer was distinguished for its efficiency and ability to handle increased parameter dimensionality effectively.

Kan, Shuwen [Fordham University]↗

Between a Map and a Data Rod

A Digital Divide has long stood between how NASA and other satellite-derived data are typically archived (time-step arrays or maps) and how hydrology and other point-time series oriented communities prefer to access those data. In essence, the desired method of data access is orthogonal to the way the data are archived. Our approach to bridging the Divide is part of a larger NASA-supported data rods project to enhance access to and use of NASA and other data by the Consortium of Universities for the Advancement of Hydrologic Science, Inc. (CUAHSI) Hydrologic Information System (HIS) and the larger hydrology community. Our main objective was to determine a way to reorganize data that is optimal for these communities. Two related objectives were to optimally reorganize data in a way that (1) is operational and fits in and leverages the existing Goddard Earth Sciences Data and Information Services Center (GES DISC) operational environment and (2) addresses the scaling up of data sets available as time series from those archived at the GES DISC to potentially include those from other Earth Observing System Data and Information System (EOSDIS) data archives. Through several prototype efforts and lessons learned, we arrived at a non-database solution that satisfied our objectivesconstraints. We describe, in this presentation, how we implemented the operational production of pre-generated data rods and, considering the tradeoffs between length of time series (or number of time steps), resources needed, and performance, how we implemented the operational production of on-the-fly (virtual) data rods. For the virtual data rods, we leveraged a number of existing resources, including the NASA Giovanni Cache and NetCDF Operators (NCO) and used data cubes processed in parallel. Our current benchmark performance for virtual generation of data rods is about a years worth of time series for hourly data (9,000 time steps) in 90 seconds. Our approach is a specific implementation of the general optimal strategy of reorganizing data to match the desired means of access. Results from our project have already significantly extended NASA data to the large and important hydrology user community that has been, heretofore, mostly unable to easily access and use NASA data.

machine learning↗

Transforming Windows from Energy Liabilities to Zero-Energy Assets: Next-Generation Solutions for Buildings

Windows have traditionally contributed to a building's HVAC load, but they can also become a source of net energy gain or even operate as zero-energy components. For heating applications, highly insulating windows can harness more solar heat than the energy lost through them, transforming windows from energy liabilities to assets. Dynamic glazings provide further benefits by regulating solar heat gain, reducing cooling loads in summer and heating demands in winter. This simulation study focuses on developing the next generation of zero-energy windows (ZEW) for residential new construction. Through annual energy simulations across climate zones 1-8, ZEW performance benchmarks were established based on current code-level buildings, and we've identified the regions where meeting ZEW standards are most achievable. This work evaluates both static and dynamic window technologies, assessing their effects on annual energy use and cost. Key findings demonstrate that ZEW performance is achievable across diverse climate zones, with specific regional requirements. Most climate zones from 3-8 can achieve ZEW with specific configurations, while some warm climates (1-2) appear challenging for ZEW implementation. Climate zones 4-6 consistently allow for zero energy window implementation, offering multiple pathways through either static or dynamic window technologies. Colder climate zones (7-8) ZEW products allow for higher SHGC values while requiring low U-values.

Yu, Lili↗

NCCDS performance model

The NASA/GSFC Network Control Center (NCC) provides communication services between ground facilities and spacecraft missions in near-earth orbit that use the Space Network. The NCC Data System (NCCDS) provides computational support and is expected to be highly utilized by the service requests needed in the future years. A performance model of the NCCDS has been developed to assess the future workload and possible enhancements. The model computes message volumes from mission request profiles and SN resource levels and generates the loads for NCCDS configurations as a function of operational scenarios and processing activities. The model has been calibrated using the results of benchmarks performed on the operational NCCDS facility and used to assess some future SN service request scenarios.

Richmond, Eric↗

Optical nanofiber testbeds for benchmarking membrane-waveguide photonic integrated circuit platforms toward on-chip quantum inertial sensing

Recent advances in cold atom interferometry with optical and magnetic atom guides have set the stage for quantum inertial sensors capable of operating in dynamic environments. In this work, we present three key innovations—evanescent-field (EF) atom guides, optical nanofiber testbeds, and membrane-waveguide photonic integrated circuit (PIC) platforms—to advance EF-guided atom interferometry. First, we demonstrate EF atom guides on optical nanofiber testbeds, which serve as performance benchmarks for our membrane-waveguide PIC platforms. Second, we achieve low-power (⁠ ~ 5 mW) guiding of freely moving, laser-cooled 133 Cs atoms in two-color, traveling-wave EF optical dipole traps at the novel, heat-efficient magic wavelengths of 793 and 937 nm (i.e., “793/937-nm EF atom guides”). Concurrently, we design and fabricate membrane-waveguide PIC platforms for these EF atom guides; in our prior work, we showed that these structures safely accommodate 4–6 times the required optical trap power under vacuum and enable dense cold atom generation via magneto-optical trapping in the vicinity of the optical wavguide for efficient loading. Third, we verify preserved atomic coherence via microwave fields and EF-coupled Doppler-free Raman beams; to our knowledge, this is the first report of coherence fringes driven by co-propagating EF-coupled Raman beams with only 150 nW of total optical power. By providing a direct comparison between optical nanofiber testbeds and membrane-waveguide PIC platforms, our results lay critical groundwork for the on-chip realization of EF-guided atom interferometry and the development of fully integrated, compact, lightweight, and low-power quantum accelerometers and gyroscopes.

Orozco, Adrian [Sandia National Laboratories (SNL-↗

Sliding mode control method having terminal convergence in finite time

An object of this invention is to provide robust nonlinear controllers for robotic operations in unstructured environments based upon a new class of closed loop sliding control methods, sometimes denoted terminal sliders, where the new class will enforce closed-loop control convergence to equilibrium in finite time. Improved performance results from the elimination of high frequency control switching previously employed for robustness to parametric uncertainties. Improved performance also results from the dependence of terminal slider stability upon the rate of change of uncertainties over the sliding surface rather than the magnitude of the uncertainty itself for robust control. Terminal sliding mode control also yields improved convergence where convergence time is finite and is to be controlled. A further object is to apply terminal sliders to robot manipulator control and benchmark performance with the traditional computed torque control method and provide for design of control parameters.

Venkataraman, Subramanian T.↗

Hardware Selection and Performance of Low-Cost Fluorometers

Access to and extensive use of fluorometric analyses is limited, despite its extensive utility in environmental transport and fate. Wide-spread application of fluorescent tracers has been limited by the prohibitive costs of research-grade equipment and logistical constraints of sampling, due to the need for high spatial resolutions and access to remote locations over long timescales. Recently, low-cost alternatives to research-grade equipment have been found to produce comparable data at a small fraction of the price for commercial equipment. Here, we prototyped and benchmarked performance of a variety of fluorometer components against commercial units, including performance as a function of tracer concentration, turbidity, and temperature, all of which are known to impact fluorometer performance. While component performance was found to be comparable to the commercial units tested, the best configuration tested obtained a functional resolution of 0.1 ppb, a working concentration range of 0.1 to >300 ppb, and a cost of USD 59.13.

42 ENGINEERING↗

Performance and Portability of a Linear Solver Across Emerging Architectures

A linear solver algorithm used by a large-scale unstructured-grid computational fluid dynamics application is examined for a broad range of familiar and emerging architectures. Efficient implementation of a linear solver is challenging on recent CPUs offering vector architectures. Vector loads and stores are essential to effectively utilize available memory bandwidth on CPUs, and maintaining performance across different CPUs can be difficult in the face of varying vector lengths offered by each. A similar challenge occurs on GPU architectures, where it is essential to have coalesced memory accesses to utilize memory bandwidth effectively. In this work, we demonstrate that restructuring a computation, and possibly data layout, with regard to architecture is essential to achieve optimal performance by establishing a performance benchmark for each target architecture in a low level language such as vector intrinsics or CUDA. In doing so, we demonstrate how a linear solver kernel can be mapped to Intel® Xeon™ and Xeon Phi™, Marvell® ThunderX2®, NEC® SX-Aurora™ TSUBASA Vector Engine, and NVIDIA® and AMD® GPUs. We further demonstrate that the required code restructuring can be achieved in higher level programming environments such as OpenACC, OCCA, and Intel® OneAPI™/SYCL, and that each generally results in optimal performance on the target architecture. Relative performance metrics for all implementations are shown, and subjective ratings for ease of implementation and optimization are suggested.

Programming models↗

Evaluation of Cathode Materials with Lithium-Metal Anodes: Baseline Performance and Protocol Standardization of Coin Cells

The collaborative evaluation of electrode materials across multiple research entities requires standardized electrochemical testing protocols to produce reliable, one-to-one comparisons between different systems of interest. Similar to the work done by Long et al. on protocol standardization for coin-cell testing with graphite anodes [J. Electrochem. Soc., 163, A2999, (2016)], here we introduce two standardized testing protocols designed to quickly evaluate important electrochemical properties of cathode materials using lithium-metal anodes. The two protocols measure kinetic and thermodynamic capacity losses, rate- and voltage-dependent cycling capacities, instabilities at high voltage and high cycling rate, and overpotentials at various states of charge. We then apply these protocols to four commercially available cathode materials to establish benchmark performance metrics that can be used to screen and evaluate new cathode materials.

25 ENERGY STORAGE↗

Emergent Behaviors and the Limits of Large Language Model Generalization

Over the past couple years, large language models (LLMs) have rapidly grown in size and performance across wide numbers of tasks, leading some to herald them as machine learning models with truly “general” capabilities. We address some of the ways that the language used to describe these models can be both deceptive and a poor framing for making progress towards understanding their capabilities, and also review arguments on to what extent these models can represent and interact with the meaning of language. Emergent behaviors in these highly complex models are unpredictable and research into understanding them is still in early phases, which limits the strength of claims of generality about them. Moreover, while these models demonstrate surprising emergent reasoning capabilities on many tasks, there are still many limitations in generalization that overall high benchmark performance can hide. We review the literature showcasing limitations of LLMs as well as techniques to mitigate or overcome these challenges, while also highlighting how fundamental problems in model evaluation may prevent true claims of generality. We conclude with a practical section on using LLMs on real world problems while appropriately evaluating and validating these tasks to prevent or mitigate the downstream impacts of incorrect results.

97 MATHEMATICS AND COMPUTING↗

Reassessing the Origins and Contemporary Relevance of ck Acceptability Parameters: Evolving Perspectives on Similarity

“Sensitivity and Uncertainty Analyses Applied to Criticality Safety Validation,” introduces sensitivity and uncertainty methods to address challenges in defining and extending areas of applicability for criticality safety validation. These areas are traditionally defined by the bounds or limits on key parameters, but establishing valid ranges and managing complex parameter variations remain challenging. NUREG/CR-6655 introduces ck and other integral indices, as well as concepts such as the completeness of benchmark coverage, to better quantify system similarities. The work proposed herein seeks to evaluate these foundational concepts to ensure that the bounds remain effective in guiding the assessment of similarity and applicability in modern applications. The concept of completeness, along with other parameters envisioned within the framework, serves as an example of the foundational ideas that have been established, though their effectiveness in practice may not be fully understood. Advancements in scripting tools, coupled with the speed and efficiency of modern computing and statistical models, now allow for faster and more thorough assessments than previously possible. These advancements also enable the identification of trends within the data, which could provide additional insight into system behavior and further broaden the scope of previously performed benchmarks. By leveraging these capabilities, we will revisit and expand the scope of these foundational methods to determine whether the necessary elements for robust similarity evaluation are already embedded, partially realized, or remain untapped.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Cloud-Based Numerical Weather Prediction for Near Real-Time Forecasting and Disaster Response

The use of cloud computing resources continues to grow within the public and private sector components of the weather enterprise as users become more familiar with cloud‐computing concepts, and competition among service providers continues to reduce costs and other barriers to entry. Cloud resources can also provide capabilities similar to high‐performance computing environments, supporting multi‐node systems required for near real‐time, regional weather predictions. Referred to as "Infrastructure as a Service", or IaaS, the use of cloud-based computing hardware in an on‐demand payment system allows for rapid deployment of a modeling system in environments lacking access to a large, supercomputing infrastructure. Use of IaaS capabilities to support regional weather prediction may be of particular interest to developing countries that have not yet established large supercomputing resources, but would otherwise benefit from a regional weather forecasting capability. Recently, collaborators from NASA Marshall Space Flight Center and Ames Research Center have developed a scripted, on‐demand capability for launching the NOAA/NWS Science and Training Resource Center (STRC) Environmental Modeling System (EMS), which includes pre‐compiled binaries of the latest version of the Weather Research and Forecasting (WRF) model. The WRF‐EMS provides scripting for downloading appropriate initial and boundary conditions from global models, along with higher‐resolution vegetation, land surface, and sea surface temperature data sets provided by the NASA Short‐term Prediction Research and Transition (SPoRT) Center. This presentation will provide an overview of the modeling system capabilities and benchmarks performed on the Amazon Elastic Compute Cloud (EC2) environment. In addition, the presentation will discuss future opportunities to deploy the system in support of weather prediction in developing countries supported by NASA's SERVIR Project, which provides capacity building activities in environmental monitoring and prediction across a growing number of regional hubs throughout the world. Capacity‐building applications that extend numerical weather prediction to developing countries are intended to provide near real‐time applications to benefit public health, safety, and economic interests, but may have a greater impact during disaster events by providing a source for local predictions of weather‐related hazards, or impacts that local weather events may have during the recovery phase.

Molthan, Andrew↗

Mechanochemically Robust LiCoO 2 with Ultrahigh Capacity and Prolonged Cyclability

Pushing intercalation-type cathode materials to their theoretical capacity often suffers from fragile Li-deficient frameworks and severe lattice strain, leading to mechanical failure issues within the crystal structure and fast capacity fading. This is particularly pronounced in layered oxide cathodes because the intrinsic nature of their structures is susceptible to structural degradation with excessive Li extraction, which remains unsolved yet despite attempts involving elemental doping and surface coating strategies. Herein, a mechanochemical strengthening strategy is developed through a gradient disordering structure to address these challenges and push the LiCoO 2 (LCO) layered cathode approaching the capacity limit (256 mAh g -1 , up to 93% of Li utilization). This innovative approach also demonstrates exceptional cyclability and rate capability, as validated in practical Ah-level pouch full cells, surpassing the current performance benchmarks. Comprehensive characterizations with multiscale X-ray, electron diffraction, and imaging techniques unveil that the gradient disordering structure notably diminishes the anisotropic lattice strain and exhibits high fatigue resistance, even under extreme delithiation states and harsh operating voltages. Consequently, this designed LCO cathode impedes the growth and propagation of particle cracks, and mitigates irreversible phase transitions. In conclusion, this work sheds light on promising directions toward next-generation high-energy-density battery materials through structural chemistry design.

36 MATERIALS SCIENCE↗

Transparent and Conductive Inorganic/Polymer‐Composite Encapsulants for Long‐Term Perovskite Solar Cells Operation

An innovative inorganic/polymer-composite encapsulation scheme comprising a polymer-based transparent conductive composite (TCC) coupled with a transparent conductive oxide is introduced to extend the lifetime of moisture-sensitive devices such as perovskite solar cells (PSC). The TCC comprises conductive silver-coated polymethyl methacrylate (Ag-PMMA) microsphere fillers protruding from a transparent non-conductive polymer matrix. TCC samples (5% Ag-PMMA by area) demonstrate high optical transparencies (%T approx. 85% in the 370–1200 nm region), low out-of-plane electrical resistivities (R < 0.2 Ω cm 2 ), and equilibrium permeabilities of less than 1 g mm/m 2 /day, all of which are well-maintained after 1,000 hours of environmental exposure. In practice, the TCC is employed between two indium zinc oxide (IZO) thin films, with the in-plane conductivity of IZO and the out-of-plane conductivity of TCC working collectively to function as a single encapsulating electrode. This encapsulation scheme is exemplified by benchmarking performances of PSCs subjected to accelerated aging conditions of quasi-maximum power point tracking under light soaking in air at 50 °C, 40% R.H. to 60% R.H. Notable improvements in operational lifetimes are observed, with a champion encapsulated PSC maintaining over 90% of its initial efficiency of 21.65% for up to 1,430 hours, compare to less than 300 hours for an unencapsulated control.

14 SOLAR ENERGY↗

blocks_3d: software for general 3d conformal blocks

We introduce the software blocks_3d for computing four-point conformal blocks of operators with arbitrary Lorentz representations in 3d CFTs. It uses Zamolodchikov-like recursion relations to numerically compute derivatives of blocks around a crossing-symmetric configuration. It is implemented as a heavily optimized, multi-threaded, C++ application. We give performance benchmarks for correlators containing scalars, fermions, and stress tensors. As an example application, we recompute bootstrap bounds on four-point functions of fermions and study whether a previously observed sharp jump can be explained using the “fake primary” effect. We conclude that the fake primary effect cannot fully explain the jump and the possible existence of a “dead-end” CFT near the jump merits further study.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Combustion dynamics of crude and upgraded Thermal DeOxygenation oils in a compression ignition engine

Thermal DeOxygenation (TDO) is a robust thermochemical conversion scheme to produce hydrocarbons from biomass feedstock. The process targets the carbohydrate fraction of biomass to yield a broad mixture of primarily aromatic hydrocarbons within a boiling point range of 348–798 K with low oxygen content (<4 wt%). The resulting materials are amenable to traditional hydrotreating and distillation processes with approximately 70% by mass in the distillate fuel range. The simple conversion scheme combined with commercially viable upgrading routes makes TDO oils a feasible alternative fuel for transportation applications. Here this paper is the first systematic treatment of fit-for-purpose testing of TDO oils in a compression ignition engine. Blends of partially upgraded TDO oils (e.g. whole oil, distilled oil, hydrotreated oil and hydrotreated-distilled oil) are prepared with certified ultra-low sulfur diesel at 5%, 10%, 15% and 20% by volume. The resulting fuels are analyzed for fuel characteristics and combustion dynamics in an instrumented single-cylinder compression ignition engine. All fuel blends at 10% blend level by volume exhibited adequate engine performance. Hydrotreating of the oils is a necessary step to meet cetane number and EPA soot emissions requirements at 20% blend volume. Combined distillation and hydrotreatment meet or exceed all fuel specification and engine performance benchmarks. Partially upgraded TDO oils are suitable fuel options for compression ignition applications.

42 ENGINEERING↗