Engineering PapersSearch

SEARCH · Engineering Papers

Results for “DATA PROCESSING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY

Time-tagging data acquisition system for testing superconducting electronics based on an RFSoC and custom analog frontend

Novel electronic devices can often be operated in a plethoraof ways, which makes testing circuits comprised of them difficult.Often, no single tool can simultaneously analyze the operatingmargins, maximum speed, and failure modes of a circuit, particularlywhen the intended behavior of subcomponents of the circuit is notstandardized. This work demonstrates a cost-effective time-domaindata acquisition system for electronic circuits that enables moreintricate verification techniques than are practical withconventional experimental setups. We use high-speeddigital-to-analog converters and real-timemulti-gigasample-per-second waveform processing to push experimentalcircuits beyond their maximum operating speed. Our customtime-tagging data capture firmware reduces memory requirements andcan be used to determine when errors occur. The firmware iscombined with a thermal-noise-limited analog frontend with50 dB of dynamic range. Compared to currentlyavailable commercial test equipment that is seven times moreexpensive, this data acquisition system was able to operate asuperconducting shift register at a nearly three-times-higher clockfrequency (200 MHz vs. 80 MHz).

Foster, Reed A. [MIT] (ORCID:0000000231002127)

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production by conducting a statistical analysis of historical wind data over a five-year period (2020-2025) from a single 1.5MW turbine manufactured by General Electric (GE) located at NLR’s Flatirons Campus, to generate an experimental test profile that was deployed on a 1.25-MW proton exchange membrane type MC250 electrolyzer system manufactured by Nel Hydrogen . [1] While the electrolyzer balance-of-plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. The historical wind data provided several metrics, however, the analysis particularly focused on the measured power output by the wind turbine. The power output time series of data for each day was categorized by total energy generation and standard deviation, and the day that represented the highest combination of these two metrics was chosen – December 25th, 2022. This process was then repeated for a moving four-hour window within this day to identify the most statistically variable period. Finally, this four-hour period was scaled by 65% to match the 1.25 MW electrolyzer. The electrolysis system controls hydrogen production by varying DC current applied to the stack, from a maximum of 3000 A to a minimum safe operation of 300 A, or 10%. Because the current – voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical wind profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1 Hz frequency. For more details on the statistical analysis process, see the presentation labeled “ Public Reference Data for Megawatt-Scale Hydrogen Electrolysis” provided with each data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}_{scaling factor}-{electrolyzer ramp rate in amperes/second} For instance, “wind-GE1.5MW_0.65-400.zip” represents the hour-long experiment using historical data from the wind-GE1.5MW turbine, scaled to 65%, with the electrolyzer power supply set to a maximum ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and wind power input. A PDF file detailing the historical wind data statistical analysis used to generate the wind profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_historical_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis [1] nelhydrogen.com/product/mc-series-electrolyser .

08 HYDROGEN

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES

CMIP7 data request: Earth system priorities and opportunities

This paper presents a comprehensive overview of the Coupled Model Intercomparison Project Phase 7 (CMIP7) request for data pertaining to Earth systems science, and provides justification for the resources needed to produce this data. Topics within the CMIP7 Earth System (CMIP7-ES) theme centre around tracking of flows of energy, carbon, water and other fluxes across domains, and constraining feedbacks between these cycles and the climate system. These topics are summarized in this paper as scientific “opportunities” describing specific model intercomparison experiments and use cases for next-generation Earth System Model (ESM) output. These opportunities were submitted by modelling groups and scientific consortia following an extended public consultation process. Contained within each opportunity are requests for groups of Climate & Forecasting (CF) variables, which are bundled into variable groups representing all data required to address the opportunities' needs. Novel opportunities in CMIP7 compared with previous phases will include running `emissions-driven' simulations that integrate carbon emissions and removal scenarios with updated representations of the global carbon cycle, expanded variable groups needed to model marine trophic interactions and biogeochemistry, and data needed to understand the risk of global tipping points, among others. The production of these variables will close key gaps and uncertainties identified during previous rounds of CMIP, and support the 7th Intergovernmental Panel on Climate Change Assessment Report (AR7). We argue that CMIP7-ES data will be broadly used by scientific, policy, governmental, industry, and other communities that rely on climate model projections for research and decision making. As an author group we also reflect on the evolution of the CMIP7-ES data request as a part of a deliberative process in support of the global CMIP program.

54 ENVIRONMENTAL SCIENCES

Using pile-up collisions as an abundant source of low-energy hadronic physics processes in ATLAS and an extraction of the jet energy resolution

During the 2015–2018 data-taking period, the Large Hadron Collider delivered proton-proton bunch crossings at a centre-of-mass energy of 13 TeV to the ATLAS experiment at a rate of roughly 30 MHz, where each bunch crossing contained an average of 34 independent inelastic proton-proton collisions. The ATLAS trigger system selected roughly 1 kHz of these bunch crossings to be recorded to disk. Offline algorithms then identify one of the recorded collisions as the collision of interest for subsequent data analysis, and the remaining collisions are referred to as pile-up. Pile-up collisions represent a trigger-unbiased dataset, which is evaluated to have an integrated luminosity of 1.33 pb -1 in 2015–2018. This is small compared with the normal trigger-based ATLAS dataset, but when combined with vertex-by-vertex jet reconstruction it provides up to 50 times more dijet events than the conventional single-jet-trigger-based approach, and does so without adding any additional cost or requirements on the trigger system, readout, or storage. The pile-up dataset is validated through comparisons with a special trigger-unbiased dataset recorded by ATLAS, and its utility is demonstrated by means of a measurement of the jet energy resolution in dijet events, where the statistical uncertainty is significantly reduced for jet transverse momenta below 65 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Advances in Modeling Capabilities for Critical Mineral Separation Technologies: A PrOMMiS Overview

This is an oral presentation at the TechConnect conference on the work developed by PrOMMiS. PrOMMiS builds on and extends capabilities developed within the Department of Energy’s (DOE) Institute for the Design of Advanced Energy Systems (IDAES), Integrated Platform, and Water Treatment Technoeconomic Assessment Platform (WaterTAP), which have been successfully leveraged by other Department of Energy research areas. The open-source toolkit facilitates validation, reproducibility, and accountability, allowing for easy extension of the framework to other systems. This talk presents an overview of the PrOMMiS capabilities, including unit model library, advances in thermophysical properties models, and capital cost libraries for simulation and optimization of mineral processing technologies. The PrOMMiS applications include (1) conceptual design and superstructure optimization for screening different process configurations and identifying promising technologies; (2) dynamic modeling and optimization to enable the creation of digital twins; (3) surrogate modeling tools to leverage data when predictive thermodynamic models are not currently available; (4) technical risk reduction via uncertainty quantification and robust optimization to identify process designs that are robust to process variability and uncertainties; and (5) deployment of uncertainty quantification tools to maximize knowledge gained from experimental campaigns, while reducing the number of experiments required

critical minerals and materials

ENVnet provides a global molecular resource of dissolved organic matter

Dissolved organic matter (DOM) is an important component of Earth's carbon cycle and one of the planet's most chemically diverse pools, yet the molecular structures of its constituents remain largely unresolved. This limitation has hindered our ability to link DOM composition to microbial processes and ecosystem function. Here we present ENVnet, a global molecular repository built from tandem mass spectrometry data collected across 13 terrestrial and aquatic environment types, including 419 newly generated samples that expand publicly available DOM metabolomics data and cover previously underrepresented environments. By computationally deconvolving chimeric mass spectra, a longstanding challenge in environmental metabolomics, we recover high-quality fragmentation data for >22,000 distinct molecular features (defined by a specific precursor mass and fragmentation pattern). Using ENVnet, we uncover conserved and environment-specific molecular patterns in DOM composition and underlying biogeochemical processes. We also use molecular features encoded in ENVnet to train predictive models of DOM persistence, allowing molecular-level assessment of microbial turnover in independent systems.

54 ENVIRONMENTAL SCIENCES

Truck Platooning Performance with ADAS and Onboard Camera Data Describing Traffic Interactions

This project was part of the Characterizing Behaviors and Capabilities for Emerging Connected and Automated Vehicle Technologies, Sensors, and Connectivity project. The National Laboratory of the Rockies partnered with Cummins Inc. to collect data from Class 8 tractor trailer combinations in platoon (cooperative adaptive cruise control) operations on public roads in southern Indiana. Data collected include J1939 CAN bus, radar, intervehicle position, and video data. The video data could not be shared in the raw form, so they were processed to extract information on the other vehicles on the road, their relative positions, and intrusion events. This information was then columnized for modeling use and further enhanced by appending road information including road type, speed limit, altitude, and grade. The test route included free-flowing traffic, highway interchanges, and construction zones, as well as low-, medium-, and high-grade sections. Individual test conditions varied by day, with advanced driver-assistance system (ADAS) features engaged or disengaged and different combined vehicle masses tested in addition to uncontrolled variables such as weather and traffic interactions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Modeling Mycorrhizal Carbon Costs in Temperate Forests: The Impacts of Functional Diversity and Global Change Factors

Mycorrhizal fungi form symbiotic relationships with most plant species, facilitating nutrient acquisition while consuming a significant fraction of the plant's photosynthetic carbon (C), which we define as the mycorrhizal C cost. Drivers of the mycorrhizal C cost, which is crucial for predicting environmental impacts on plant productivity, remain under-explored and difficult to quantify. Ecosystem models that incorporate mycorrhizae can offer insights into mycorrhizal C cost dynamics, but their predictions have rarely been validated against empirical data. Here, in this study, we used the Myco-CORPSE model, which explicitly simulates mycorrhizal processes alongside soil carbon and nitrogen cycling, to investigate the drivers of mycorrhizal C cost in temperate forests. Applying this model to over 1,800 forest inventory plots across the eastern United States, we found that the simulations matched published data, showing higher C allocation to ectomycorrhizal (ECM) fungi (16.0% of net primary production (NPP)) compared to arbuscular mycorrhizal (AM) fungi (5.8% of NPP). Further analysis showed that mixed forests, co-dominated by both AM and ECM trees, allocated less C to mycorrhizal fungi compared to forests dominated by either AM or ECM fungi alone, due to complementary nutrient acquisition strategies. Elevated Nitrogen (N) deposition and higher temperatures reduce mycorrhizal C costs, favoring AM strategies. Conversely, elevated CO 2 (eCO 2 ) increased plant N demand and mycorrhizal C costs, favoring ECM strategies that access organic N sources. These findings underscore the critical role of mycorrhizal functional diversity in plant nutrient acquisition and C dynamics, providing new insights into how mycorrhizal symbioses respond to global change.

Shao, Siya [Dartmouth College, Hanover, NH (United

Recommended Nuclear Structure and Decay Data for A=206 Isobars

Here, evaluated nuclear structure and decay data for all nuclei with mass number A=206 ( 206 Pt, 206 Au, 206 Hg, 206 Tl, 206 Pb, 206 Bi, 206 Po, 206 At, 206 Rn, 206 Fr, 206 Ra and 206 Ac), are presented. All available experimental data are compiled and evaluated, and best values for level and γ-ray energies, quantum numbers, lifetimes, γ-ray intensities and transition probabilities, as well as other nuclear properties, are recommended. Inconsistencies and discrepancies that exist in the literature are discussed. A number of computer codes (https://www-nds.iaea.org/public/ensdf_pgm/) developed by members of the NSDD network were used during the evaluation process. This work supersedes the earlier evaluation by F.G. Kondev (2008Ko21), published in Nuclear Data Sheets 109, 1527 (2008).

Kondev, F. G. [Argonne National Laboratory (ANL),

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (

VISIONARY: Virtual Intelligence System for Optimizing Novel Analytical Research Yields

VISIONARY is an AI system that accelerates energy materials discovery by automatically generating hypotheses about structure-property relationships. It analyzes patterns in materials data, identifies promising correlations, and proposes testable scientific hypotheses without human intervention. By streamlining this reasoning process, VISIONARY helps researchers efficiently identify candidate materials with desired properties, significantly speeding up the materials development pipeline for energy applications. During the project, we developed a standalone application. The application uses a combination of papers provided by the user and data collected from FutureHouse’s dataset to build an understanding of the background that the user wants to explore for the hypothesis.

36 MATERIALS SCIENCE

CMIP7 data request: land and land ice priorities and opportunities

The Land and Land Ice Theme in the Coupled Model Intercomparison Project Phase 7 (CMIP7) represents the current understanding of physical processes in land surface ecosystems, hydrology, cryosphere, and their physical interactions with other Earth system components. Simulations from Earth system models (ESMs) could provide crucial information for assessing planetary safety, such as critical tipping elements, and be used to inform climate risks for improving climate impact assessments and policy decisions. This paper presents a collaborative effort to identify scientific opportunities in the Land and Land Ice Theme of the CMIP7 Data Request. The proposed opportunities build upon advances in ESMs, including new freshwater system and land ice processes being included in CMIP7, as well as the scientific community's demand for high-frequency and sub-grid-scale land surface outputs. In total, 25 variable groups that contain 716 variables have been identified to be potentially available to the broad scientific audience for performing analysis in land–atmosphere coupling, hydrological processes and freshwater systems, glacier and ice sheet mass balance and their influence on the sea levels, land use, and plant phenology. Key reflections from this data request effort include advocacy for closer engagement between the user community and modeling groups, reduction in the technical barriers to tracking existing parameters and defining new variables, and more streamlined variable management. These will be essential to enhance the usability and reliability of CMIP7 outputs for climate and Earth system research and applications to a broad audience that relies on the CMIP7 endeavor.

Li, Yue [Univ. of California, Los Angeles, CA (Uni

Optimization and quantification of silver( II ) for mediated electrochemical oxidation applications

Mediated electrochemical oxidation (MEO) is a low-temperature, low-pressure, aqueous mineralization process used to treat organic waste. A powerful metal oxidant is used as a mediator in an acidic solution. Although Ce and Co are thoroughly studied mediators, Ag is a preferred choice because of the higher efficiency rates of mineralization observed with this system. Importantly, the quantification methodology and spectroscopic characteristics of the Ag(II) ion must be obtained. In this study, we determined molar extinction coefficients of the primary absorption band associated with the Ag(II) ion in 2–9 M HNO 3 solution. The optimization of Ag(II) electrooxidation was also determined by altering parameters such as HNO 3 concentration, mediator concentration, and temperature. The optimization studies and extinction coefficient data provide parameters for implementation of Ag as a suitable mediator for MEO processing of organic waste.

Schrage, Briana R. [Oak Ridge National Laboratory

A critical review on additive manufacturing of refractory alloys from a data analytics perspective- beyond nickel-based superalloys

Refractory alloys (RAs) are promising materials due to their exceptional physicochemical properties, but most research remains at the laboratory scale. For broader adoption, advancements in manufacturing are essential. Because their high stability makes conventional methods like machining and casting difficult, additive manufacturing (AM) is emerging as an effective approach for fabricating refractory alloy components. However, AM's repeated non-equilibrium thermal cycles introduce undesired features (e.g. defects, anisotropic microstructures, and residual stresses), which are magnified due to RAs’ unique properties. This paper comprehensively reviews the state-of-the-art methods of AM for refractory alloys. It explores data analytics techniques to establish design rules based on multi-fidelity experimental and computational methods. Furthermore, it investigates integrated, collaborative efforts to harmonise standalone databases, information, knowledge, and predictive models at multi-physics, multi-stage, and multi-scale. Unlike the existing literature that focuses primarily on material systems or process fundamentals, this work provides an integrated perspective on AM of refractory alloys from a data analytics standpoint, highlighting the roles of integrated computational materials engineering (ICME), verification, validation, and uncertainty quantification (VV&UQ), and digital twin-driven qualification in overcoming data scarcity and accelerating rapid qualification.

Additive manufacturing