Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

The NASA ACTIVATE Mission

The NASA Aerosol Cloud Meteorology Interactions over the Western Atlantic Experiment (ACTIVATE) conducted 162 joint flights with two aircraft over the northwest Atlantic to study aerosol–cloud interactions (ACIs), which represent the largest uncertainty in estimating total anthropogenic radiative forcing. The combination of a high-flying King Air and low-flying HU-25 Falcon, equipped with remote sensing and in situ instruments, characterized trace gases, aerosol particles, clouds, and meteorological variables with data collected nearly simultaneously below, within, and above marine boundary layer (MBL) clouds. Flights spanning warm and cold seasons across 3 years (2020–22) provided a broad range of conditions associated with aerosol particles, cloud properties (including particle size and phase), and meteorology, ideally suited for robust ACI calculations and assessing how well models simulate a wide range of MBL clouds from stratiform to cumulus. ACTIVATE data suggest that drivers of cloud droplet number concentration N d , including aerosol particles and MBL dynamics, vary between winter and summer months with a stronger potential to convert aerosol particles into cloud droplets in winter. Models of varying complexity not only highlight some skills in simulating winter and summer cloud types but also identify challenges that still need to be addressed such as treatment of turbulence, wet scavenging, and mesoscale organization. Remote sensing advances range from new retrieval methods for N d , cloud phase classification, vertically resolved aerosol and cloud condensation nuclei number concentration, and ocean surface wind speed. This work describes these scientific and technological advances along with efforts in outreach and open data science.

aerosol indirect effect↗

Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)

This data release provides all data and code used in the paper " "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)" to model stream temperature, evaluate, and assess results. The associated manuscript explores current open questions in prediction in ungauged and unmonitored basins concerning top-down versus bottom-up approaches, tradeoffs between data available and input requirements, and the appropriate representation of catchment attributes as inputs to deep learning models. Modeling was done primarily with long short-term memory (LSTM) models, and stream site coverage spans 1362 locations across the conterminous United States. The data is organized into these items items:Code repository and data for the paper " "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)".Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code: - data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- error_analysis_attribute_and_groundwater_dir.zip - workflows for the extended error analysis by stream attribute and groundwater influenceData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2024streamdata, author = {Jared Willard and Fabio Ciulla and Helen Weierbach and Vipin Kumar and Charuleka Varadharajan}, title = {Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models"}, year = {2024}, doi = {10.15485/2448016}, publisher = {ESS-DIVE Repository}, url = {https://doi.org/10.15485/2448016}}MLA: Willard, Jared, et al. Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models". 2024. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Canopy tree mortality and crown exposure data from the Amacayacu Forest Dynamics Plot, Northwestern Amazon

Data on the mortality of 984 canopy trees, their crown exposure to light (relative to total crown area), growth deviations (relative to conspecifics), tree size, and species’ wood density collected between 2013 and 2019 in 18 ha of the Amacayacu Forest Dynamics Plot, Northwestern Amazon. This dataset contains a single CSV data file. Variable definitions: 1. Species: [character] species identification 2. Family: [character] family of the species 3. tag: [character] unique consecutive for the tree 4. status: [character] status of the tree in the third census (ALIVE or DEAD) 5. wsg: [numeric]: species’ wood density (g cm-3) 6. growth_r1: [numeric]: annual growth rate of the tree between the first and second census (cm y-1) 7. growth_r2: [numeric]: annual growth rate of the tree between the second and third census (cm y-1) 8. gt1: [numeric] modulus transformed growth rate with a lambda of 0.4 for the annual growth rate of the tree between the first and second census (cm y-1) 9. gt2: [numeric]: modulus transformed growth rate with a lambda of 0.4 for the annual growth rate of the tree between the second and third census (cm y-1) 10. rgr1: [numeric] relative growth rate between the first and second census (cm y-1) 11. rgr2: [numeric] relative growth rate between the second and third census (cm y-1) 12. sa_gt: [numeric] species-adjusted modulus transformed growth rate 13. sa_rgr: [numeric] species-adjusted relative growth rate 14. gr_n: [numeric] number of individuals of the species used to calculate the mean and species modulus transformed growth rate 15. dbh_flight: [numeric] diameter at breast height (1.3 m) estimated at the time of the drone flight (cm) 16. eca: [numeric] exposed crown area calculated as the area of the crown polygon delineated in the orthomosaic (m2) 17. total_ca: [numeric] total crown area estimated from a crown area model (m2) 18. rcel: [numeric] relative crown exposure to light (m2) 19. time_flight_census3: [numeric] time in years from the drone flight date to the third census for that tree (yr)

54 ENVIRONMENTAL SCIENCES↗

Probabilistic data fusion and physics-informed machine learning: A new paradigm for modeling under uncertainty, and its application to accelerating the discovery of new materials

In this report we summarize the work conducted by PI Perdikaris and his group under this Early Career project DE–SC0019116 during the period of 09/01/2018 – 08/31/2023. The central aim of the work was to introduce a new paradigm for scientific data analysis that can seamlessly synthesize rigorous mathematical modeling with data of variable fidelity (e.g., measurements at multiple scales/resolutions or predictions of variable fidelity models) and multiple modalities (e.g., images, time–series, or scattered measurements). The setting we are interested in involves complex systems that are partially observed and whose dynamical behavior could be hard to model or totally unknown. The inherent uncertainty associated with this setting necessitates a departure from the classical deterministic realm of modeling and scientific computation, and, consequently, our main building blocks can no longer be crisp deterministic numbers and governing laws, but instead we must operate with probabilistic models.

97 MATHEMATICS AND COMPUTING↗

Observations and Lessons Learned in Residential Roofing Integrated Photovoltaics

The data file includes field observations of residential roofing integrated photovoltaics installation that happened between July 2021 and June 2022 in California. There are 21 observations - 2 in the re-roofing category and 19 in the new construction category. The file includes two time and motion forms (Re-Roofing & New Construction) we used to capture the time it takes for each activity in the rooftop solar and electrical installation process. You will see the duration for detailed steps along with activities that are included for total installation time for the paper. There are three parts to the time and motion form - Part 1 has project and crew characteristics, product information, and inspection and permitting data and can be filled before the installation. Part 2 has the actual installation time stamps and related notes. Part 3 has the crew information and can be filled on-site or off-site, but this section is optional. The re-roofing form has pre-solar activities and gutters and vents section in Part 2 of the form, whereas new construction form includes a section for capturing the time it took for rough wiring. The re-roofing form does not include rough wiring but instead includes the final electrical wiring process. These are the key differences in both the forms. Each day/step/installation activity needs to include the crew break time as they happen and there is space available to capture the crew breaks duration. In Table 1 below, the different terms used in the time and motion form are defined and their corresponding unit of measurement stated. We also highlight the optional sections. The data is captured in total mins, total hours, person mins, person hours. You can use this form or make edits to the form to recreate the study or make your own observations. The data herein was reviewed but may not be comprehensive. NREL invites questions and inputs to improve the data, including to: Correct erroneous information Fill in missing/updated information Clarifications on data and variables Updated information may be submitted to Sushmita Jena at sushmita.jena@nlr.gov.

14 SOLAR ENERGY↗

Overcoming Communications Outages in Inverter Downtime Analysis: Preprint

Inverters are often reported to be the highest-impact failure point in PV systems. This importance is belied by the simplistic assumptions about inverter downtime losses used in industrial energy modeling. Energy models often assume the extent of inverter-related losses is limited to 1% of annual production from scheduled inverter preventative maintenance. In reality, inverter-related production losses are much more variable, although data from large-scale surveys of fielded systems are rare. Here we present a method of detecting inverter downtime events and estimating the associated lost production using inverter- and meter-level power data. Because communications outages are of similar frequency to true production outages, the method pays particular attention to distinguishing communications outages from true production outages. The results of applying the method at fleet-scale are presented and discussed.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Unraveling the Wrinkle in Time-Variable Sources with Lunes and Synthetic Seismic Data

In this report, we describe how to estimate the time-variable components of the seismic moment tensor and compare these estimates to the more conventional analysis that incorporates an assumption of the source time function (STF) across all components of the seismic moment tensor. The advantage of our method is that we are able to independently estimate the time-evolution of each component of the seismic moment tensor, which may help to resolve the complex source phenomena associated with buried explosions. By performing an eigen decomposition of the time-evolving seismic moment tensor components, we are able to plot the seismic mechanism as a trajectory on a lune diagram. This technique enables interpretation of the seismic mechanism as a function of time, as opposed to the more conventional analysis which assumes that the seismic mechanism is time invariant. Finally, we describe the differences between the seismic moment and the seismic moment rate STFs, how to implement each one in inversion schemes, and the relative strengths/weaknesses of each. Our key take-away is that we are able to distinguish nearly-overlapping sources with highly different mechanisms, such as an explosion immediately following an earthquake, by estimating moment rate from seismic data through a STF-invariant inversion for the full time-variable moment tensor.

58 GEOSCIENCES↗

Physical discovery in representation learning via conditioning on prior knowledge

Recent advances in electron, scanning probe, optical, and chemical imaging and spectroscopy yield bespoke data sets containing the information of structure and functionality of complex systems. In many cases, the resulting data sets are underpinned by low-dimensional simple representations encoding the factors of variability within the data. The representation learning methods seek to discover these factors of variability, ideally further connecting them with relevant physical mechanisms. However, generally, the task of identifying the latent variables corresponding to actual physical mechanisms is extremely complex. Here, we present an empirical study of an approach based on conditioning the data on the known (continuous) physical parameters and systematically compare it with the previously introduced approach based on the invariant variational autoencoders. The conditional variational autoencoder (cVAE) approach does not rely on the existence of the invariant transforms and hence allows for much greater flexibility and applicability. Interestingly, cVAE allows for limited extrapolation outside of the original domain of the conditional variable. However, this extrapolation is limited compared to the cases when true physical mechanisms are known, and the physical factor of variability can be disentangled in full. We further show that introducing the known conditioning results in the simplification of the latent distribution if the conditioning vector is correlated with the factor of variability in the data, thus allowing us to separate relevant physical factors. We initially demonstrate this approach using 1D and 2D examples on a synthetic data set and then extend it to the analysis of experimental data on ferroelectric domain dynamics visualized via piezoresponse force microscopy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Updates to Relevance Vector Machine: Multiclass Classification, Variable Selection, and Proof-of-Concept Application to Safeguards Fresh Fuel Verification using List-Mode Neutron Collar Data

To expand the capabilities of safeguards authorities to verify the integrity of fresh fuel assemblies, Oak Ridge National Laboratory has retrofit the existing electronics of the JCC-71 uranium neutron coincidence collar, which contains 18 3 He neutron detectors and an external 241 AmLi(α, n) neutron interrogation source arranged to surround a fresh nuclear fuel assembly. The new electronics system allows analysts to record list-mode neutron multiplicity data in addition to the singles and doubles rates that are currently measured. Based on previous proof-of-concept research, analysis of these new data will identify off-normal fuel configurations in an assembly and characterize or localize the specific partial fuel defects. The purpose of this report it to document the analysis algorithm development and then to demonstrate its capability for the safeguards verification of fresh fuel assemblies using list mode neutron collar data. To analyze the complex list-mode data collected with the upgraded uranium neutron collar, multivariate classification algorithms are being developed using a novel classification method, the relevance vector machine. This approach may be applied to multiclass problems to estimate the probability that test data belongs to one of many possible classes of data. In addition, our method identifies the most useful variables/channels for making predictions, which illuminates the basis for the model’s predictions, and this interpretability is largely unique among data analytics methods. Variable selection occurs during model training and parameter tuning and does not need any external hyperparameter tuning routines. Finally, we apply the modified relevance vector machine to a simulated dataset of list-mode neutron collar data generated with the radiation transport code MCNP. The method can correctly identify off-normal fuel configurations, categorize the data according to four fuel defect scenarios, and rank the channels in the data according to prediction utility. For nuclear safeguards applications, it is concluded that this method has the potential to increase the sensitivity and reliability to detect missing fuel rods from a standard 17 x 17 Pressurized Water Reactor (PWR) fresh fuel assembly. Within this analysis, “off-normal” (i.e., missing fuel rods) were correctly classified in 17 simulated test scenarios with one quarter (25%) of the fresh fuel rods missing using a training data set of 58 simulated measurements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Benchmarking Variables for Checkpointing in HPC Applications

Checkpoint/Restart (C/R) is a widely used fault tolerance mechanism in converged systems of cloud, edge, and HPC. However, users often rely on their experience to determine which variables to checkpoint, as there is currently no benchmark that can provide a reference. This can result in checkpointing redundant or even incorrect variables. To address this issue, we propose a benchmark suite that includes critical variables for checkpointing, which have been manually identified, and a method for identifying those critical variables, with 20 representative HPC applications. Our method involves analyzing data dependency between variables to identify critical variables analytically. We verify the identified variables' correctness with a widely used C/R library FTI by an ablation study. With our benchmark suite and data dependency analysis, HPC practitioners now have a reference for identifying checkpointing variables and better knowledge of what kind of variables to checkpoint.

Fu, Xiang↗

Data for Aboveground rather than belowground productivity drives variability in Miscanthus x giganteus net primary productivity

This dataset contains the data used for the publication “Aboveground rather than belowground productivity drives variability in Miscanthus x giganteus net primary productivity”. This dataset contains Miscanthus x giganteus biomass, carbon, and nitrogen tissue data for aboveground and belowground plant parts collected in 2021 for three different sites in Iowa with three different nitrogen application rates. Data at the Iowa sites were collected via biometric hand harvesting, belowground excavations, and soil coring both in-clump and beside-clump. Data were collected at two collection timepoints to calculate the contributions of belowground parts to Miscanthus x giganteus net primary productivity. This dataset also includes Miscanthus x giganteus and Switchgrass soil coring and excavation data collected in 2012 at the University of Illinois Urbana Champaign Energy Farm.

Belowground Biomass↗

Landmark-embedded Gaussian process with applications for functional data modeling

In practice, we often need to infer the value of a target variable from functional observation data. A challenge in this task is that the relationship between the functional data and the target variable is very complex: the target variable not only influences the shape but also the location of the functional data. In addition, due to the uncertainties in the environment, the relationship is probabilistic, that is, for a given fixed target variable value, we still see variations in the shape and location of the functional data. To address this challenge, we present a landmark-embedded Gaussian process model that describes the relationship between the functional data and the target variable. A unique feature of the model is that landmark information is embedded in the Gaussian process model so that both the shape and location information of the functional data are considered simultaneously in a unified manner. Gibbs-Metropolis-Hasting algorithm is used for model parameters estimation and target variable inference. The performance of the proposed framework is evaluated by extensive numerical studies and a case study of nano-sensor calibration.

42 ENGINEERING↗

Data for Intra- and inter-annual variability of nitrification in the rhizosphere of field-grown bioenergy sorghum

These data were collected in 2018 and 2019 at the University of Illinois Energy Farm (N 40.063607, W 88.206926). During each growing season, bulk and rhizosphere soil were collected from replicate Sorghum bicolor nitrogen use efficiency trial plots at three separate time points (approximately July 1, August 1, and September 1). We measured soil moisture, pH, soil nitrate and ammonium, potential nitrification, potential denitrification, and extracted and sequenced the V4 region of the 16S rRNA gene for microbial community analysis. All microbial sequence data is archived in the National Center for Biotechnology Information’s (NCBI) Sequence Read Archive (accession number SRP326979, project number PRJNA741261).

bioenergy↗

Predicting the heat release variability of Li-ion cells under thermal runaway with few or no calorimetry data

Accurate measurement of the variability of thermal runaway behavior of lithium-ion cells is critical for designing safe battery systems. However, experimentally determining such variability is challenging, expensive, and time-consuming. Here, we utilize a transfer learning approach to accurately estimate the variability of heat output during thermal runaway using only ejected mass measurements and cell metadata, leveraging 139 calorimetry measurements on commercial lithium-ion cells available from the open-access Battery Failure Databank. We show that the distribution of heat output, including outliers, can be predicted accurately and with high confidence for new cell types using just 0 to 5 calorimetry measurements by leveraging behaviors learned from the Battery Failure Databank. Fractional heat ejection from the positive vent, cell body, and negative vent are also accurately predicted. We demonstrate that by using low cost and fast measurements, we can predict the variability in thermal behaviors of cells, thus accelerating critical safety characterization efforts.

25 ENERGY STORAGE↗

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis↗

Model code and data: biomass allocation adjustments induced by elevated CO2 and warming in a C3 brackish marsh, 2017-2022, Maryland

This dataset and R script accompany the published paper Bruns et al. (2024) in Geophysical Research Letters. The data are from the first six years of a field manipulation of whole-ecosystem warming and elevated CO2 experiment (Salt Marsh Accretion Response to Temperature eXperiment, or SMARTX) in the Smithsonian's Global Change Research Wetland (GCReW), a brackish, microtidal wetland site on a subestuary of the Chesapeake Bay. These data were generated to understand how warming and elevated CO2 interact to structure ecosystem-level responses to global change, particularly in terms of carbon sequestration. The dataset covers 2017-2022 and includes peak annual above ground biomass, annual belowground fine root productivity, and porewater NH4 for each experimental plot. The overall experiment is replicated in two locations on the marsh, a lower elevation zone dominated the C3 sedge S. Americanus and a higher elevation plot dominated by the C4 species. This paper and its data release is only for the C3 plot. Variable descriptions for data file is available in variable_descriptions.pdf. The R script Bruns_et_al_2024_GRL_make_figures.Rmd contains model code and other scripts used to generate all paper figures.

54 ENVIRONMENTAL SCIENCES↗

Comment on Comment on “Anomalous structural recovery in the near glass transition range in a polymer glass: Data revisited in light of temperature variability in vacuum oven‐based experiments”*

Abstract Cangialosi, Alegría, and Colmenero have made a comment on a paper of ours [Polym. Eng. Sci. 2022:1–13], in which we discussed the concern that the enthalpy recovery data reported by Cangialosi and co‐workers [Phys. Rev. Lett. 2013;111(9):095701] for polystyrene aged up to 15 K below glass transition temperature was anomalous and contradicted existing experimental results from the literature over a similar range of aging conditions. Their response shifts the focus away from the raised questions about their experimental results and attempts to invalidate the data that we cited in support of our argument. Here we respond to the comment and add additional analysis that suggests the structural recovery response of glassy materials exhibits smooth behavior over the full range of measurements, up to 7 or 8 logarithmic decades. We do this by referring to the work on the intrinsic isotherm down‐jump and memory responses of poly(vinyl acetate) between 40°C and 15°C over six logarithmic decades by Kovacs [Fortsch. Hochpolym. Fo. 1963;3(1/2):394–508], the small‐strain tensile creep of poly(vinyl chloride) quenched from 90°C to 40°C (approximately 40°C below T g ) over seven logarithmic decades from Struik [Polym Eng. Sci., 1977;17:165–173], and the volume recovery behavior for aging times up to 3 months in a temperature range between 95°C and −50°C by Greiner and Schwarzl [Rheol. Acta. 1984;23(4):378–395]. We also add discussion that an isothermal aging procedure using a vacuum oven is highly vulnerable to temperature errors due to the problem of good temperature control when the heat transfer mechanism is primarily radiative.

Jin, Shuang↗