Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data reduction pipelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Selection of a Pair of Experiments to Optimally Reduce Uncertainty in Targeted Nuclear Data

We propose a novel process to select a pair of differential and integral experiments that best reduce uncertainties in targeted 239 ⁢Pu nuclear data while compressing the current nuclear data pipeline from 20 to 3 years. 239⁢ Pu nuclear data are poorly understood for neutrons in the intermediate energy range due to sparsity and uncertainty in historical experiments. New experiments targeting this range will enable better understanding of these nuclear data, but choosing the ideal experiments to conduct is challenging. Beginning with a prior distribution represented by samples of nuclear data generated from theory, generalized least squares adjustments are made to incorporate data from historical experiments. To quantify potential uncertainty reduction obtainable from a pair of candidate experiments, we compute the D-optimality criterion of the posterior covariance of intermediate energy range nuclear data compared to the equivalent covariance after additional adjustment to the pair of candidate experiments. Repeating the process for each of many candidate pairs facilitates the final selection. Results support 63⁢ Cu total cross section measurements for differential experiments and alumina and alumina/graphite configurations for integral experiments. This analysis enables choosing differential and integral experiments to be executed concurrently while shortening decision times relative to the current nuclear data pipeline.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

Central Stars of Planetary Nebulae in the LMC

In FUSE cycle 2's program B001 we studied Central Stars of Planetary Nebulae (CSPN) in the Large Magellanic Could. All FUSE observations have been successfully completed and have been reduced, analyzed and published. The analysis and the results are summarized below. The FUSE data were reduced using the latest available version of the FUSE calibration pipeline (CALFUSE v2.2.2). The flux of these LMC post-AGB objects is at the threshold of FUSE's sensitivity, and thus special care in the background subtraction was needed during the reduction. Because of their faintness, the targets required many orbit-long exposures, each of which typically had low (target) count-rates. Each calibrated extracted sequence was checked for unacceptable count-rate variations (a sign of detector drift), misplaced extraction windows, and other anomalies. All the good calibrated exposures were combined using FUSE pipeline routines. The default FUSE pipeline attempts to model the background measured off-target and subtracts it from the target spectrum. We found that, for these faint objects, the background appeared to be over-estimated by this method, particularly at shorter wavelengths (i.e., < 1000 A). We therefore tried two other reductions. In the first method, subtraction of the measured background is turned off and and the background is taken to be the model scattered-light scaled by the exposure time. In the second one, the first few steps of the pipeline were run on the individual exposures (correcting for effects unique to each exposure such as Doppler shift, grating motions, etc). Then the photon lists from the individual exposures were combined, and the remaining steps of the pipeline run on the combined file. Thus, more total counts for both the target and background allowed for a better extraction.

Bianchi, Luciana

Joint US-Japan Observations with the Infrared Space Observatory (ISO): Deep Surveys and Observations of High-Z Objects

Several important milestones were passed during the past year of our ISO observing program: (1) Our first ISO data were successfully obtained. ISOCAM data were taken for our primary deep field target in the 'Lockman Hole'. Thirteen hours of integration (taken over 4 contiguous orbits) were obtained in the LW2 filter of a 3 ft x 3 ft region centered on the position of minimum HI column density in the Lockman Hole. The data were obtained in microscanning mode. This is the deepest integration attempted to date (by almost a factor of 4 in time) with ISOCAM. (2) The deep survey data obtained for the Lockman Hole were received by the Japanese P.I. (Yoshi Taniguchi) in early December, 1996 (following release of the improved pipeline formatted data from Vilspa), and a copy was forwarded to Hawaii shortly thereafter. These data were processed independently by the Japan and Hawaii groups during the latter part of December 1996, and early January, 1997. The Hawaii group made use of the U.S. ISO data center at IPAC/Caltech in Pasadena to carry out their data reduction, while the Japanese group used a copy of the ISOCAM data analysis package made available to them through an agreement with the head of the ISOCAM team, Catherine Cesarsky. (3) Results of our LW2 Deep Survey in the Lockman Hole were first reported at the ISO Workshop "Taking ISO to the Limits: Exploring the Faintest Sources in the Infrared" held at the ISO Science Operations Center in Villafranca, Spain (VILSPA) on 3-4 February, 1997. Yoshi Taniguchi gave an invited presentation summarizing the results of the U.S.-Japan team, and Dave Sanders gave an invited talk summarizing the results of the Workshop at the conclusion of the two day meeting. The text of the talks by Taniguchi and Sanders are included in the printed Workshop Proceedings, and are published in full on the Web. By several independent accounts, the U.S.-Japan Deep Survey results were one of the highlights of the Workshop; these data showed conclusively that the ISOCAM S/N continues to decrease as the square root of time for periods as long as 13 hours.

Sanders, David B.

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

Computer-Aided Parallelizer and Optimizer

The Computer-Aided Parallelizer and Optimizer (CAPO) automates the insertion of compiler directives (see figure) to facilitate parallel processing on Shared Memory Parallel (SMP) machines. While CAPO currently is integrated seamlessly into CAPTools (developed at the University of Greenwich, now marketed as ParaWise), CAPO was independently developed at Ames Research Center as one of the components for the Legacy Code Modernization (LCM) project. The current version takes serial FORTRAN programs, performs interprocedural data dependence analysis, and generates OpenMP directives. Due to the widely supported OpenMP standard, the generated OpenMP codes have the potential to run on a wide range of SMP machines. CAPO relies on accurate interprocedural data dependence information currently provided by CAPTools. Compiler directives are generated through identification of parallel loops in the outermost level, construction of parallel regions around parallel loops and optimization of parallel regions, and insertion of directives with automatic identification of private, reduction, induction, and shared variables. Attempts also have been made to identify potential pipeline parallelism (implemented with point-to-point synchronization). Although directives are generated automatically, user interaction with the tool is still important for producing good parallel codes. A comprehensive graphical user interface is included for users to interact with the parallelization process.

Jin, Haoqiang

Machine Learning Techniques for Data Reduction of Climate Applications

Scientists conduct large-scale simulations to compute derived quantities-of-interest (QoI) from primary data. Often, QoI are linked to specific features, regions, or time intervals, such that data can be adaptively reduced without compromising the integrity of QoI. For many spatiotemporal applications, these QoI are binary in nature and represent presence or absence of a physical phenomenon. We present a pipelined compression approach that first uses neural-network-based techniques to derive regions where QoI are highly likely to be present. Then, we employ a Guaranteed Autoencoder (GAE) to compress data with differential error bounds. GAE uses QoI information to apply low-error compression to only these regions. This results in overall high compression ratios while still achieving downstream goals of simulation or data collections. Experimental results are presented for climate data generated from the E3SM Simulation model for downstream quantities such as tropical cyclone and atmospheric river detection and tracking. These results show that our approach is superior to comparable methods in the literature.

Li, Xiao [University of Florida]

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

Large reductions in Permian Basin methane intensity shown in multi-year comparison of aerially-visible methane emissions

Spanning the US states of Texas and New Mexico, the Permian Basin has been a hotspot of methane emissions from oil and natural gas activity 1–4, although studies disagree over the magnitude of these emissions. The most comprehensive measurement campaigns published were conducted in 2019 1–5 .There have been large changes in the energy industry since then, including in the prices of oil and gas, both state and federal regulatory environments, investor and activist pressure over methane emissions, and the adoption of new technologies and policies by energy operators. Understanding how any or all of these might influence methane emissions is important for policy makers, oil and gas operators, and other stakeholders. We characterize the time evolution of Permian Basin methane emissions using a series of comprehensive aerial surveys conducted every year from 2020-2023 and compare them to the 2019 results cited above. To maintain comparability, all the data sets are from surveys using Insight M point source methane sensing technology. The scope of these surveys expanded over time: from 33-46% of wells, oil production, and gas production in 2020 to 60-65% in 2021, to 84% of wells and over 90% of both oil and gas production in 2023. These surveys by Insight M also include hundreds of gas processing plants and compressor stations as well as 1000s of km of gathering and transmission pipelines. Considering only the aerially detected portion of emissions (typically the majority of the total in such surveys 4), we find reductions of more than 70% in methane emissions intensity compared to the 2019 New Mexico-only Insight M survey, with variation depending on the year 3,4. Notably, although sources below 100 kg/hr contributed less than 10% of aerially measured emissions the 2019 New Mexico survey 3, these smaller sources constitute a larger proportion of total aerially measured emissions (although not the majority) in 2020-2023. Production facilities and gathering pipelines are responsible for the larges shares of total emissions, followed by compressor stations and gas processing plants. Permian methane emissions were also measured in a comprehensive 2019 Permian-wide survey by the Carbon Mapper team 2. That analysis led to a lower total emissions estimate at the time 4. These new Insight M-based emission rates are still roughly 30-70% lower than the aerially measured portion of the 2019 Carbon Mapper-based estimates 4. Further work is needed to harmonize these surveys in space and time to create the most intercomparable numbers possible 5. Additional analysis is needed to compare our findings to the more spatially constrained 2020, 2021, and 2023 Carbon Mapper surveys in the Permian 4,6. The evidence is strong from these two survey teams that emissions intensity has declined significantly since 2019. Reasons for this trend are currently unclear but point to possible success of emissions control programs. Future work investigating frequency, source, and operator-specific intensities could provide insights into the causes of this promising trend.

methane, oil and gas, data science, remote sensing

Personalized and uncertainty-aware coronary hemodynamics simulations: From Bayesian estimation to improved multi-fidelity uncertainty quantification

Non-invasive simulations of coronary hemodynamics have improved clinical risk stratification and treatment outcomes for coronary artery disease, compared to relying on anatomical imaging alone. However, simulations typically use empirical approaches to distribute total coronary flow amongst the arteries in the coronary tree, which ignores patient variability, the presence of disease, and other clinical factors. Further, uncertainty in the clinical data often remains unaccounted for in the modeling pipeline. We present an end-to-end uncertainty-aware pipeline to (1) personalize coronary flow simulations by incorporating vessel-specific coronary flows as well as cardiac function; and (2) predict clinical and biomechanical quantities of interest with improved precision, while accounting for uncertainty in the clinical data. We assimilate patient-specific measurements of myocardial blood flow from clinical CT myocardial perfusion imaging to estimate branch-specific coronary artery flows. Simulated noise in the clinical data is used to estimate the joint posterior distributions of the model parameters using adaptive Markov Chain Monte Carlo sampling. Additionally, the posterior predictive distribution for the relevant quantities of interest is determined using a new approach combining multi-fidelity Monte Carlo estimation with non-linear, data-driven dimensionality reduction. This leads to improved correlations between high- and low-fidelity model outputs. Our framework accurately recapitulates clinically measured cardiac function as well as branch-specific coronary flows under measurement noise uncertainty. We observe substantial reductions in confidence intervals for estimated quantities of interest compared to single-fidelity Monte Carlo estimation and state-of-the-art multi-fidelity Monte Carlo methods. This holds especially true for quantities of interest that showed limited correlation between the low- and high-fidelity model predictions. In addition, the proposed multi-fidelity Monte Carlo estimators are significantly cheaper to compute than traditional estimators, under a specified confidence level or variance. The proposed pipeline for personalized and uncertainty-aware predictions of coronary hemodynamics is based on routine clinical measurements and recently developed techniques for CT myocardial perfusion imaging. The proposed pipeline offers significant improvements in precision and reduction in computational cost.

Bayesian parameter estimation

Latent Twins

Over the past decade, scientific machine learning has transformed the development of mathematical and computational frameworks for analyzing, modeling, and predicting complex systems. From inverse problems to numerical partial differential equations (PDEs), dynamical systems, and model reduction, these advances have pushed the boundaries of what can be simulated. Yet they have often progressed in parallel, with representation learning and algorithmic solution methods evolving largely as separate pipelines. With Latent Twins, we propose a unifying mathematical framework that creates a hidden surrogate in latent space for the underlying equations. Whereas digital twins mirror physical systems in the digital world, Latent Twins mirror mathematical systems in a learned latent space governed by operators. Through this lens, classical modeling, inversion, model reduction, and operator approximation all emerge as special cases of a single principle. We establish the fundamental approximation properties of Latent Twins for both ordinary differential equations (ODEs) and PDEs and demonstrate the framework across three representative settings: (i) canonical ODEs, capturing diverse dynamical regimes; (ii) a PDE benchmark using the shallow-water equations, contrasting Latent Twin simulations with deep operator network and forecasts with a four-dimensional variational method baseline; and (iii) a challenging real-data geopotential reanalysis dataset, reconstructing and forecasting from sparse, noisy observations. Latent Twins provide a compact, interpretable surrogate for solution operators that evaluate across arbitrary time gaps in a single-shot, while remaining compatible with scientific pipelines such as assimilation, control, and uncertainty quantification. Looking forward, this framework offers scalable, theory-grounded surrogates that bridge data-driven representation learning and classical scientific modeling across disciplines.

Latent Twins

ROSAT Science Data Center

This report provides a summary of the Smithsonian Astrophysical Observatory (SAO) ROSAT SCIENCE DATA CENTER (RSDC) activities for the recent years of our contract. Details have already been reported in the monthly reports. The SAO was responsible for the High Resolution Imager (HRI) detector on ROSAT. We also provided and supported the HRI standard analysis software used in the pipeline processing (SASS). Working with our colleagues at the Max Planck in Garching Germany (MPE), we fixed bugs and provided enhancements. The last major effort in this area was the port from VMS/VAX to VMS/ALPHA architecture. In 1998, a timing bug was found in the HRI standard processing system which degraded the positional accuracy because events accessed incorrect aspect solutions. The bug was fixed and we developed off-line correction routines and provided them to the community. The Post Reduction Off-line Software (PROS) package was developed by SAO and runs in the IRAF environment. Although in recent years PROS was not a contractual responsibility of the RSDC, we continued to maintain the system and provided new capabilities such as the ability to deal with simulated AXAF data in preparation for the NASA call for proposals for Chandra. Our most recent activities in this area included the debugging necessary for newer versions of IRAF which broke some of our software. At SAO we have an operating version of PROS and hope to release a patch even though almost all functionality that was lost was subsequently recovered via an IRAF patch (i.e. most of our problems were caused by an IRAF bug).

Murray, Stephen

Outline of a fast hardware implementation of Winograd's DFT algorithm

The main characteristics of the discrete Fourier transform (DFT) algorithm considered by Winograd (1976) is a significant reduction in the number of multiplications. Its primary disadvantage is a higher structural complexity. It is, therefore, difficult to translate the reduced number of multiplications into faster execution of the DFT by means of a software implementation of the algorithm. For this reason, a hardware implementation is considered in the current study, taking into account a design based on the algorithm prescription discussed by Zohar (1979). The hardware implementation of a FORTRAN subroutine is proposed, giving attention to a pipelining scheme in which 5 consecutive data batches are being operated on simultaneously, each batch undergoing one of 5 processing phases.

Zohar, S.

Optimization of foreground moment deprojection for semi-blind CMB polarization reconstruction

Abstract Upcoming Cosmic Microwave Background (CMB) experiments, aimed at measuring primordial CMB polarization B-modes, require exquisite control of instrumental systematics and Galactic foreground contamination. Blind minimum-variance techniques, like the Needlet Internal Linear Combination (NILC), have proven effective in reconstructing the CMB polarization signal and mitigating foregrounds and systematics across diverse sky models without suffering from foreground mismodelling errors. Still, residual foreground contamination from NILC may bias the recovered CMB polarization at large angular scales when confronted with the most complex foreground scenarios.By adding constraints to NILC to deproject statistical moments of the Galactic emission, the Constrained Moment ILC (cMILC) method has been demonstrated to further enhance foreground subtraction, albeit with an associated increase in overall noise variance. Faced with this trade-off between foreground bias reduction and overall variance minimization, there is still no recipe on which moments to deproject and which are better suited for blind variance minimization. To address this, we introduce the optimized cMILC (ocMILC) pipeline, which performs full automated optimization of the required number and set of foreground moments to deproject, pivot parameter values, and deprojection coefficients across the sky and angular scales, depending on the actual sky complexity, available frequency coverage, and experiment sensitivity. The optimal number of moments for deprojection, before paying significant noise penalty, is determined through a data diagnosis inspired by the Generalized NILC (GNILC) method.Validated on B-mode simulations of thePICOspace mission concept with four challenging foreground models, ocMILC exhibits lower Galactic foreground contamination compared to NILC and cMILC at all angular scales, with limited noise penalty. This multi-layer optimization enables the ocMILC pipeline to achieve unbiased posteriors of the tensor-to-scalar ratio, regardless of foreground complexity.

Astronomy & Astrophysics

Utility-Scale Solar, 2024 Edition: Empirical Trends in Deployment, Technology, Cost, Performance, PPA Pricing, and Value in the United States [Slides]

Berkeley Lab’s “Utility-Scale Solar, 2024 Edition” presents analysis of empirical plant-level data from the U.S. fleet of ground-mounted photovoltaic (PV), PV+battery, and concentrating solar-thermal power (CSP) plants with capacities exceeding 5 MWAC (PV plants of 5 MWAC or less, including residential rooftop systems, are covered separately in Berkeley Lab’s companion annual report, Tracking the Sun). Key findings from this year’s report include: -18.5 GWAC of new utility-scale PV capacity came online in 2023, bringing cumulative installed capacity to more than 80.2 GWAC across 47 states. Installed costs continued to fall in 2023. Relative to 2022, capacity-weighted averages decreased by 8% to -$\$1.43$/WAC (or $\$1.08$/WDC). Costs, based on a 7.1 GWAC sample of 76 plants completed in 2023, have fallen by 75% (averaging 10% annually) since 2010. Plant-level capacity factors vary widely, from 6% to 36% (on an AC basis), with a sample median of 24%. -Levelized cost of energy (LCOE) of new 2023 projects increased slightly to $\$46$/MWh prior to the application of tax credits but continued to fall to $\$31$/MWh when accounting for federal incentives. PPA prices have largely followed the decline in solar’s LCOE over time, but newly signed longer-term PPA prices have increased since 2021, to an average of $\$35$/MWh (levelized, in 2023 dollars). -Solar’s average energy and capacity value (i.e., ability to offset costs of other power generation sources) across the U.S. was $\$45$/MWh in 2023. Solar’s average market value was lowest in CAISO ($\$27$/MWh), the market with the greatest solar generation share, and highest in ERCOT ($\$67$/MWh). -Newer solar projects had greater market value in 2023 than their generation costs, yielding $\$1.1$ billion in benefits. Projects built in 2022 delivered on average $\$15$/MWh more market value than their costs in 2023. -Solar’s combined value from wholesale electricity markets, public health and climate damage reduction were greater than generation costs and incentives, yielding $\$13.7$ billion in net benefits in 2023. We estimate U.S. health benefits of $\$24$/MWh and reduced global climate damages of $\$101$/MWh. -Adding battery storage is one way to increase the value of solar. Deployment of 52 new PV+battery hybrid plants set a record with 5.3 GW installed in 2023. Our public data file tracks metadata and PPA prices from more than 100 PV+battery hybrid projects that are already online or that have secured offtake arrangements. -Looking ahead, a massive pipeline of at least 1,085 GW of solar capacity dominates the nation’s interconnection queues at the end of 2023. Nearly 571 GW, or 53%, of that total was paired with a battery – in CAISO it was a staggering 98%. Historically only 10% of the requested solar capacity is built. -For more information, and to explore related interactive data visualizations, go to utilityscalesolar.lbl.gov.

14 SOLAR ENERGY

Importance of Higher Fidelity Model Geometries during Optimization of Critical Experiments

PARADIGM, PARallel Approach of Differential and InteGral Measurements, is a cross-collaborative effort at Los Alamos National Laboratory between nuclear data theorists, differential and integral experimenters, as well as machine learning statisticians to tackle uncertainties in the intermediate region of 239 Pu. In essence, the idea behind PARADIGM is to remove the linear conceptualization of the nuclear data pipeline, shown in Figure 1, and replace it with a far more parallelized approach. The novel approach leverages machine learning to guide which differential measurements and integral experiments will result in the largest decrease in uncertain ties for a nuclide reaction pair in a given energy range. The concept builds off earlier work, EUCLID, which focused on the fast region of 239 Pu. The practical benefit of having evaluation, differential measurement, and integral experiment personnel in collaboration with machine learning is to represent the entire nuclear data in one snapshot. This enable large reduction in the time to deliver improved nuclear data, which using the PARADIGM approach could be done in 3 years. A general outline of PARADIGM and specific topics are available in other papers. The discussion here will pertain directly to the integral experiment design. More specifically, the process of taking a rough design and transforming it into a finalized neutronic model will be discussed.

97 MATHEMATICS AND COMPUTING

Assessment of lnternational Space Station (ISS) Lithium-ion Battery Thermal Runaway (TR)

This task was developed in the wake of the Boeing 787 Dreamliner lithium-ion battery TR incidents of January 2013 and January 2014. The Electrical Power Technical Discipline Team supported the Dreamliner investigations and has followed up by applying lessons learned to conduct an introspective evaluation of NASA's risk of similar incidents in its own lithium-ion battery deployments. This activity has demonstrated that historically NASA, like Boeing and others in the aerospace industry, has emphasized the prevention of TR in a single cell within the battery (e.g., cell screening) but has not considered TR severity-reducing measures in the event of a single-cell TR event. center dotIn the recent update of the battery safety standard (JSC 20793) to address this paradigm shift, the NASA community included requirements for assessing TR severity and identifying simple, low-cost severity reduction measures. This task will serve as a pathfinder for meeting those requirements and will specifically look at a number of different lithium-ion batteries currently in the design pipeline within the ISS Program batteries that, should they fail in a Dreamliner-like incident, could result in catastrophic consequences. This test is an abuse test to understand the heat transfer properties of the cell and ORU in thermal runaway, with radiant barriers in place in a flight like test in on orbit conditions. This includes studying the heat flow and distribution in the ORU. This data will be used to validate the thermal runaway analysis. This test does not cover the ambient pressure case. center dotThere is no pass/ fail criteria for this test.

Graika, Jason

Modeling performance of data collection systems for high-energy physics

Exponential increases in scientific experimental data are outpacing silicon technology progress, necessitating heterogeneous computing systems—particularly those utilizing machine learning (ML)—to meet future scientific computing demands. The growing importance and complexity of heterogeneous computing systems require systematic modeling to understand and predict the effective roles for ML. We present a model that addresses this need by framing the key aspects of data collection pipelines and constraints and combining them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by development vectors including ML, parallelization, advancing CMOS, and neuromorphic computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, our model allows alternative data collection systems to be rigorously compared. We apply this model to the Compact Muon Solenoid experiment and its planned high luminosity-large hadron collider upgrade, evaluating novel technologies for the data acquisition system (DAQ), including ML-based filtering and parallelized software. The results demonstrate that improvements to early DAQ stages significantly reduce resources required later, with a power reduction of 60% and increased relevant data retrieval per unit power (from 0.065 to 0.31 samples/kJ). However, we predict that further advances will be required in order to meet overall power and cost constraints for the DAQ.

Olin-Ammentorp, Wilkie (ORCID:0000000224729862)