Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data cleaning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS

UBW (USLCI-Brightway2) [SWR-25-169]

Life cycle inventory (LCI) data are critical for robust life cycle assessment (LCA), yet many widely used datasets such as the U.S. Life Cycle Inventory (USLCI) are not natively compatible with advanced modeling frameworks like Brightway2. This work presents an automated pipeline to transform USLCI data into a fully functional Brightway2 project. The workflow performs systematic data cleaning, resolves duplicate process and exchange identifiers, and applies allocation to multi-output processes. Technosphere and biosphere flows are harmonized through unit conversions and a bridge mapping to the biosphere3 database, with comprehensive logging of missing flows and cutoff issues. The resulting Brightway2 database is validated using matrix diagnostics to ensure consistency of the technosphere, and is benchmarked via life cycle impact assessment (LCIA) methods such as ReCiPe and IPCC GWP. Outputs include reproducible CSV exports of corrected processes, elementary flows, characterization factors, and LCIA results, alongside backup utilities for project sharing. This pipeline lowers barriers for integrating USLCI data into open-source LCA workflows, enabling reproducible, validated LCA inventories within the Brightway 2 framework.

Ghosh, Tapajyoti [National Laboratory of the Rocki

Characterizing peak electricity demand for U.S. households: an assessment of end-use loads and demand factors

Understanding household peak electricity demand is critical to evaluate the technical need for electrical infrastructure upgrades. This study characterizes peak loads for existing and new equipment using metered data from a convenience sample of 11,940 U.S. dwellings from four sources, including 911 from two sources with end-use metering. After standardized data cleaning and labeling, we derived descriptive statistics for key metrics, such as maximum demand and demand factors, and developed predictive models relating 60- to 15-min demand for the National Electrical Code (NEC). Mean 15-min maximum demand was 9.7 kW (median 9.0 kW; IQR 7.0–11.5 kW, 95% CI 9.6–9.8 kW), indicating spare capacity in 98% of homes with hypothetical 100 A panels. Maximum demand increased with floor area and number of high-demand loads. Dwelling maximum demand was driven by higher-power, longer-duration heating appliances and vehicle charging, while most user-operated appliances contributed little. Demand factors are used to account for how most devices contribute less than their rated power to maximum demand. Existing load mean demand factors (28%; median 10%; IQR 0–58%; CI 28–29%) were higher than those for new loads (21%; median 7%; IQR 0–35%; CI 20–21%), because new loads changed the timing and magnitude of maximum demand. New high-demand loads had higher than average demand factors (40–60%). Whole dwelling demand factors support the NEC's 40% assumption, but they challenge its conservative 100% treatment of new HVAC. We propose a data-driven 50% demand factor for new equipment, which would align with metered data, improve affordability, and modernize electrical codes.

Appliances

Signal processing and spectral modeling for the BeEST experiment

The Beryllium Electron capture in Superconducting Tunnel junctions (BeEST) experiment searches for evidence of heavy neutrino mass eigenstates in the nuclear electron capture decay of 7 Be by precisely measuring the recoil energy of the 7 Li daughter. In Phase III, the BeEST experiment has been scaled from a singl superconducting tunnel junction (STJ) sensor to a 36-pixel array to increase sensitivity and mitigate gamma-induced backgrounds. Phase III also uses a new continuous data acquisition system that greatly increases the flexibility for signal processing and data cleaning. Here, we have developed procedures for signal processing and spectral fitting that are sufficiently robust to be automated for large datasets. Furthermore, this article presents the optimized procedures before unblinding the majority of the Phase III dataset to search for physics beyond the standard model.

6 ≤ A ≤ 19

Data & Code from Phoenix CPPP Phase 2 Analysis

This data and code package supports the analysis presented in “Beyond Surface Cooling: Comprehensive Field Assessment of Reflective Pavement Thermal Performance in Phoenix, Arizona” and provides fully reproducible workflows for evaluating the thermal performance of cool pavement treatments in a hot urban environment. The dataset integrates multi-modal field measurements collected across residential and nonresidential settings, including mobile air temperature traverses, stationary air temperature monitoring, residential mean radiant temperature (MRT) measurements, subsurface temperature profiles, and controlled testbed observations. The data package contains raw and processed datasets in comma-separated value (CSV) format, accompanying metadata files describing site characteristics and measurement protocols, and R scripts (.R files) used for data cleaning, time synchronization, spatial and temporal matching, quality control filtering, statistical comparison, and figure generation. All analyses were conducted using R (version ≥ 4.2.0) with commonly available packages (e.g., tidyverse, lubridate, data.table, ggplot2). No proprietary software is required to reproduce results. Field campaigns were designed to quantify the effects of high-reflectance pavement coatings on surface temperature, near-surface air temperature, subsurface heat propagation, and radiative heat exposure. Temporal alignment procedures include standardized timestamp conversion and nearest-neighbor matching of high-frequency sensor measurements to stop-based metadata within defined tolerance windows to ensure comparability across instruments. The workflows generate summary statistics, treatment–control contrasts, depth-dependent thermal gradients, and time-series visualizations used in the associated publication. By integrating mobile, stationary, radiative, and subsurface measurements within a unified and transparent processing framework, this package enables comprehensive evaluation of cool pavement performance across multiple thermal exposure pathways and supports reuse in future urban heat mitigation and climate resilience studies.

AIR TEMPERATURE

Technical Track on Biomass Carbon Removal and Storage (BiCRS): Mapping bioresources, phase 1 - Consistency check comparing Mission Innovation’s Data Visualization Tool for Bioresources and the Clean Energy Ministerial Biofuture Initiative Global Biomass data accessible via the US Department of Energy’s Bioenergy Knowledge Discovery Framework (KDF)

The Mission Innovation (MI) Carbon Dioxide Removal (CDR) Mission, Technical Track on Biomass Carbon Dioxide Removal and Storage (BiCRS), has produced a biomass resource database for its members. In parallel, Oak Ridge National Laboratory (ORNL) developed the International Feedstock Reporting data portal—herein referred to as the CEM Biofuture-KDF data—on behalf of the Clean Energy Ministerial Biofuture Initiative (CEM Biofuture), as a specific task under Biofuture’s 2024–25 Action Plan. This work was conducted at the request of CEM Biofuture and funded by the U.S. Department of Energy in support of that initiative, and it is hosted within DOE’s Knowledge Discovery Framework (KDF).

09 BIOMASS FUELS

Automated point dendrometer, soil moisture and temperature, and meteorological variables datasets, Oct 2024 – Nov 2025, G.A. Pearson Natural Area, Flagstaff, AZ, USA

This data package includes parsed, cleaned, and calibrated data from 48 TOMST automated point dendrometers, 48 TOMST 15 cm soil moisture sensors, and 12 TOMST 30 cm soil moisture sensors. The point dendrometers were cleaned with the “dendRoAnalyst” package in RStudio. The soil sensors were cleaned and calibrated for volumetric water content (VWC) with the “myClim” package in RStudio using the soil texture of the site (sandy clay loam). Additionally, this data package also includes raw data from 2 METER weather stations. Dendrometers and soil sensors have both their sensor ID, as well as the ID for the specific tree they were instrumented on at the G.A. Pearson Natural Area (GPNA) site and their experimental group. The purpose of these data is to understand how ponderosa pine trees in restored (thinned and burned) vs. unrestored (no treatment) areas are responding to drought and seasonal precipitation. These data use radial growth and soil moisture data to answer the following question: how are active season length, growth on different time scales (weekly, monthly, seasonally, and annually), growth during dry periods and after precipitation events, and environmental and biological drivers of radial growth different between restored versus unrestored areas?

Air temperature

Review of Presentation by Carter Wolf

The presentation was informative on the different techniques used to study the electronic dynamics of materials on the femtosecond scale, such as ARPES and TR ARPES. The presentation also discussed in depth code analysis, and the procedure used to remove noise from noisy data via machine learning procedures. The presentation first focused on an overview of the ARPES technique, and how TR-ARPES differs from it. The presentation then explored computational techniques used for fitting ARPES and TR-ARPES data, using the pyARPES framework. After that, the presentation discussed machine learning approaches for denoising noisy data such as Noise2Noise and Noise2Self. The main issue of denoising existing ARPES data and the necessity of machine learning approaches due to the lack of clean ARPES data were very clearly articulated as a major part of the project. The implementation, training and refinement and optimization were clearly detailed, along with a full documentation of attempts that worked and attempts that did not, thus thoroughly and clearly showing the research process. Experimental techniques regarding ARPES were also briefly documented.

36 MATERIALS SCIENCE

ChemEcho v1.0

ChemEcho is a tool that converts tandem mass spectra into embeddings used to build machine learning (ML) models with fully explainable predictions. It provides an API for transforming raw tandem mass spectral data into embeddings, along with functions for training and validating ML models. Additionally, it includes utilities for retrieving and cleaning training data. ChemEcho is broadly applicable in ML pipelines that use tandem mass spectra for a variety of tasks, such as chemical classification or bioactivity mining. While there are existing methods to generate embeddings from fragmentation data, ChemEcho's approach ensures that predictions remain interpretable, enabling experts to evaluate results and generate hypotheses about the underlying data.

Harwood, Thomas [Lawrence Berkeley National Labora

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]

Mapping and Synthesis of International Biomass Supply Assessments

This report, Mapping and Synthesis of International Biomass Supply Assessments (or Global Biomass Resource Assessment) is the first step in a long-term process to assemble data from around the globe into a virtual repository that can be updated and provide user-friendly access to the data. The Clean Energy Ministerial (CEM) Biofuture Platform Initiative recommended that research be completed to “address the need for internationally accepted benchmarks quantifying sustainable biomass feedstock supplies.” To act upon the CEM Biofuture recommendation, in 2024, the U.S. Department of Energy (DOE) commissioned Oak Ridge National Laboratory (ORNL) to prepare this report as the primary deliverable for a one-year assignment to assemble data into a citable form that could help resolve the persistent question presented related to bioenergy policy, “Is there enough sustainable biomass?” In response to that query, this report includes information received by August 2024 from national CEM representatives, collaborators, and public sources on current and future sustainable biomass supplies in 62 nations,and subsequently documents (a) the approach used by ORNL to analyze and categorize the information received in a manner that enables aggregation and comparability; and (b) recommendations for next steps and guidelines to help others update and harmonize future assessments of global sustainable biomass supplies.

09 BIOMASS FUELS

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis

Moltensaltpropnet

MoltenSaltPropnet is a physics-informed machine learning framework that aims to predict the thermophysical properties of molten fluoride and chloride salt mixtures, which are crucial for the design and safety of Generation IV molten salt reactors. The code processes data from the Molten-Salt Thermal Properties Database (MSTDB-TP) and the Janz compendium, converting critically evaluated correlations into fast, differentiable surrogate models for density, viscosity, thermal conductivity, and heat capacity across 448 distinct salt systems. The implementation consists of several key components: 1. Data Curation: The code parses and cleans the raw data, normalizing elemental mole fractions and extracting relevant regression coefficients for various thermophysical properties. 2. Feature Engineering: It generates fixed-length numerical descriptors that encapsulate the composition and temperature, incorporating polynomial interaction terms and dimensionality-reduction techniques to optimize model performance. 3. Coefficient Learning: Four different machine learning architectures are employed: a deep residual network (ResNet), a Kolmogorov–Arnold network (KAN), a sparsity-inducing neural network (SNN), and classical regression models. Each model learns to predict coefficients that define the temperature-dependent correlations for the thermophysical properties. 4. Property Reconstruction: The predicted coefficients are used to compute temperature-dependent property values, ensuring positivity and monotonic trends through a composite loss function that enforces physical constraints. 5. User Interface: An open-source web application enables users to filter the database, train task-specific models, and visualize the results, allowing for rapid exploration of candidate salt mixtures. MoltenSaltPropnet bridges the gap between limited experimental data and high-fidelity reactor simulations, providing a powerful tool for researchers in the field of molten salt reactors and advanced nuclear energy systems.

Retamales, Mauricio Eduardo Tano [Idaho National L

Addressing Issues with Working Memory in Video Object Segmentation

Contemporary state-of-the-art video object segmentation (VOS) models compare incoming unannotated images to a history of image-mask relations via affinity or cross-attention to predict object masks. We refer to the internal memory state of the initial image-mask pair and past image-masks as a working memory buffer. While the current state of the art models perform very well on clean video data, their reliance on a working memory of previous frames leaves room for error. Affinity-based algorithms include the inductive bias that there is temporal continuity between consecutive frames. To account for inconsistent camera views of the desired object, working memory models need an algorithmic modification that regulates the memory updates and avoid writing irrelevant frames into working memory. A simple algorithmic change is proposed that can be applied to any existing working memory-based VOS model to improve performance on inconsistent views, such as sudden camera cuts, frame interjections, and extreme context changes. The resulting model performances show significant improvement on video data with these frame interjections over the same model without the algorithmic addition. Our contribution is a simple decision function that determines whether working memory should be updated based on the detection of sudden, extreme changes and the assumption that the object is no longer in frame. By implementing algorithmic changes, such as this, we can increase the real-world applicability of current VOS models.

97 MATHEMATICS AND COMPUTING

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory

Magnetic edge fields in UTe 2 near zero background fields

Chiral superconductors are theorized to exhibit spontaneous edge currents. Here, in this study, we found magnetic fields at the edges of UTe 2 , a candidate odd-parity chiral superconductor, that seem to agree with predictions for a chiral order parameter. However, we did not detect the chiral domains that would be expected, and recent polar Kerr and muon spin relaxation data in nominally clean samples argue against chiral superconductivity. Our results show that hidden sources of magnetism must be carefully ruled out when using spontaneous edge currents to identify chiral superconductivity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Queued Up: 2025 Edition – Characteristics of Power Plants Seeking Transmission Interconnection As of the End of 2024 [Slides]

Electric transmission system operators (ISOs, RTOs, or utilities) require proposed power plants seeking to connect to the transmission grid to undergo a series of impact studies before they can be built. This process establishes what new transmission equipment or upgrades may be needed before a project can connect to the system and assigns the costs of that equipment. The lists of projects in this process are known as “interconnection queues”. In collaboration with interconnection.fyi, Berkeley Lab compiled, aggregated, and cleaned interconnection queue data from >50 transmission grid operators (7 ISO/RTOs and 49 non-ISO balancing areas), which collectively represent ~97% of currently installed U.S. electric generating capacity. The dataset includes requests submitted to queues through the end of 2024, and only includes requests seeking to connect to the transmission grid (not distribution-connected or behind-the-meter projects). The files below include both a PDF report and an Excel data file. The PDF report analyzes interconnection data and metrics through the end of 2024. The Excel data file includes (a) the full project-level interconnection queue dataset through 2024, (b) a codebook (data dictionary) describing each data field, and (c) 35 additional tabs featuring tables summarizing a range of interconnection metrics. Key highlights from the Queued Up: 2025 Edition (featuring data through 2024) include: • As of the end of 2024, there were ~10,300 projects actively seeking grid interconnection in the U.S., representing 1,400 GW of generation and approximately 890 GW of storage. • Historic withdrawal rates alongside relatively fewer new requests resulted in a 12% decrease in total active queue volume compared to the prior year. • Active natural gas capacity (136 GW, +72% year-over-year) increased in 2024, while solar (956 GW, -12%), storage (890 GW, -13%), and wind (271 GW, -26%) capacity decreased. • 408 GW of capacity already has a draft or executed interconnection agreement (IA) but has not yet reached commercial operations. • The time projects spend in queues before reaching COD is increasing. For the regions with available data, the median duration from IR to COD has doubled from <2 years for projects built in 2000-2007 to over 4 years for those built in 2018-2024. • Ultimately, most of this proposed capacity will not be built. Only 13% of capacity that submitted interconnection requests from 2000-2019 had reached commercial operations by the end of 2024; 77% of that capacity had been withdrawn and 10% was still active. • FERC Order 2023 and various other reforms are being implemented. These are important measures to reduce interconnection bottlenecks and enhance grid system reliability, but it is too early to measure and assess their full impact. • New additions for the 2025 edition include: (a) additional detail on data processing and gaps; (b) updates on interconnection reforms; (c) new analysis on interconnection agreements, and more.

24 POWER TRANSMISSION AND DISTRIBUTION