Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enrichment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Using DNA affinity purification sequencing (DAP-seq) to identify in vitro binding sites of potential Novosphingobium aromaticivorans DSM12444 transcription factors

Genome-wide binding sites of 44 putative transcription factors (TFs) from Novosphingobium aromaticivorans DSM12444 were analyzed using DNA affinity purification sequencing. We report that 32 of these TFs have at least one area of enrichment. These data will help better understand aromatic metabolism and other features of N. aromaticivorans biology.

DAP-seq↗

Using DNA Affinity Purification sequencing (DAP-seq) to identify in vitro binding sites of transcription factors potentially involved in aromatic degradation

The genome-wide binding sites of 44 transcription factors from the aromatic metabolizing Alphaproteobacterium Novosphingobium aromaticivorans were identified using DNA Affinity Purification sequencing (DAP-seq). We report 32 of these transcription factors have at least one area of enrichment. These data will be valuable for better understanding of aromatic metabolism.

aromatic metabolism↗

A measurement of the antiproton flux in the cosmic rays

A balloon-borne instrument has been used to detect cosmic-ray antiprotons. These are identified topologically by the appearance of annihilation prongs in a thick lead-plate spark chamber. The initial recording of the data is enriched in potential antimatter events by a selective trigger. After a small subtraction for background, 14 identified antiprotons yield a flux of 1.7 plus or minus 0.00005 antiproton/(sq m ster sec MeV) between 130 and 320 MeV at the top of the atmosphere. When combined with higher energy antiproton flux measurements, this result indicates that the antiprotons have a spectrum whose shape is the same as that of the protons, but with a magnitude reduced by a factor of 1/3000.

Buffington, A.↗

The relationship between geology and geochemistry in the Undarum/Spumans/Balmer region of the moon

Based on regional maps of orbital geochemical variables, major geochemical heterogeneities in the Undarum/Sumans/Balmer region of the moon are found which are usually associated with distinct geological features, and which reveal a north/south dichotomy. The mapped mare and plains units are often the site of heterogeneity, and the plains units in the Balmer basin are probably composed of volcanic material, possibly KREEP-enriched basalt. Data suggest that the northeastern part of the region may contain a minor spinel component. Major mare units of the region are found to have two distinct compositions, possibly indicating two different source regions for the basalts deposited there, but more likely indicating contamination with different amounts of highland debris.

Clark, P. E.↗

Modeling Calcium Loss from Bones During Space Flight

Calcium loss from bones during space flight creates a risk for astronauts who travel into space, and may prohibit space flights to other planets. The problem of calcium loss during space flight has been studied using animal models, bed rest (as a ground-based model), and humans in-flight. In-flight studies have typically documented bone loss by comparing bone mass before and after flight. To identify changes in metabolism leading to bone loss, we have performed kinetic studies using stable isotopes of calcium. Oral (Ca-43) and intravenous (Ca-46) tracers were administered to subjects (n=3), three-times before flight, once in-flight (after 110 days), and three times post-flight (on landing day, and 9 days and 3 months after flight). Samples of blood, saliva, urine, and feces were collected for up to 5 days after isotope administration, and were analyzed for tracer enrichment. Tracer data in tissues were analyzed using a compartmental model for calcium metabolism and the WinSAAM software. The model was used to: account for carryover of tracer between studies, fit data for all studies using the minimal number of changes between studies, and calculate calcium absorption, excretion, bone calcium deposition and bone calcium resorption. Results showed that fractional absorption decreased by 50% during flight and that bone resorption and urinary excretion increased by 50%. Results were supported by changes in biochemical markers of bone metabolism. Inflight bone loss of approximately 250 mg Ca/d resulted from decreased calcium absorption combined with increased bone resorption and excretion. Further studies will assess the time course of these changes during flight, and the effectiveness of countermeasures to mitigate flight-induced bone loss. The overall goal is to enable human travel beyond low-Earth orbit, and to allow for better understanding and treatment of bone diseases on Earth.

Wastney, Meryl E.↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗

Genome-resolved metagenomics reveals role of iron metabolism in drought-induced rhizosphere microbiome dynamics

Recent studies have demonstrated that drought leads to dramatic, highly conserved shifts in the root microbiome. At present, the molecular mechanisms underlying these responses remain largely uncharacterized. Here we employ genome-resolved metagenomics and comparative genomics to demonstrate that carbohydrate and secondary metabolite transport functionalities are overrepresented within drought-enriched taxa. These data also reveal that bacterial iron transport and metabolism functionality is highly correlated with drought enrichment. Using time-series root RNA-Seq data, we demonstrate that iron homeostasis within the root is impacted by drought stress, and that loss of a plant phytosiderophore iron transporter impacts microbial community composition, leading to significant increases in the drought-enriched lineage, Actinobacteria. Finally, we show that exogenous application of iron disrupts the drought-induced enrichment of Actinobacteria, as well as their improvement in host phenotype during drought stress. Collectively, our findings implicate iron metabolism in the root microbiome’s response to drought and may inform efforts to improve plant drought tolerance to increase food security.

59 BASIC BIOLOGICAL SCIENCES↗

Two decades of fumigation data from the Soybean Free Air Concentration Enrichment facility

Abstract The Soybean Free Air Concentration Enrichment (SoyFACE) facility is the longest running open-air carbon dioxide and ozone enrichment facility in the world. For over two decades, soybean, maize, and other crops have been exposed to the elevated carbon dioxide and ozone concentrations anticipated for late this century. The facility, located in East Central Illinois, USA, exposes crops to different atmospheric concentrations in replicated octagonal ~280 m 2 Free Air Concentration Enrichment (FACE) treatment plots. Each FACE plot is paired with an untreated control (ambient) plot. The experiment provides important ground truth data for predicting future crop productivity. Fumigation data from SoyFACE were collected every four seconds throughout each growing season for over two decades. Here, we organize, quality control, and collate 20 years of data to facilitate trend analysis and crop modeling efforts. This paper provides the rationale for and a description of the SoyFACE experiments, along with a summary of the fumigation data and collation process, weather and ambient data collection procedures, and explanations of air pollution metrics and calculations.

60 APPLIED LIFE SCIENCES↗

R -Matrix Analysis and Statistical Properties of Dysprosium Isotopes in the Neutron Energy Ranges Up To A Few Kev

In support of the Nuclear Criticality Safety Program, a set of evaluated resonance parameters was generated for seven dysprosium isotopes in the neutron energy range from thermal up to a few keV. The evaluation methodology used the Reich-Moore approximation to fit, with the R-matrix code SAMMY, the high-resolution capture and transmission measurements on natural and enriched samples recently performed at the Rensselaer Polytechnic Institute Gaerttner LINear ACcelerator facility. Additional transmission data measured on enriched samples by Liou at the Columbia University Nevis synchrocyclotron in the mid-seventies were used to gauge the neutron widths above 15 eV. Thermal constants such as absorption and (in)coherent scattering cross sections and corresponding scattering lengths were calibrated to the National Institute of Standards and Technology’s compilation except for 161,164 Dy isotopes.

07 ISOTOPE AND RADIATION SOURCES↗

A Review of Candidates for a Validation Data Set for High-Assay Low-Enrichment Uranium Fuels

Many advanced reactor concept designs rely on high-assay low-enriched uranium (HALEU) fuel, enriched up to approximately 19.75% 235 U by weight. Efforts are underway by the US government to increase HALEU production in the United States to meet anticipated needs. However, very few data exist for validation of computational models that include HALEU, beyond a few fresh fuel benchmark specifications in the International Reactor Physics Experiment Evaluation Project. Nevertheless, there are other data with potential value available for developing into quality benchmarks for use in data- and software-validation efforts. This paper reviews the available evaluated HALEU fuel benchmarks and some of the potentially relevant benchmarks for fresh highly enriched uranium. It then introduces experimental data for HALEU fuel irradiated at Idaho National Laboratory, from relatively recent irradiation programs at the Advanced Test Reactor. Such data should be evaluated and, if valuable, collected into detailed benchmark specifications to meet the needs of HALEU-based reactor designers.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

GeneLab Analysis Working Group Kick-Off Meeting

Goals to achieve for GeneLab AWG - GL vision - Review of GeneLab AWG charter Timeline and milestones for 2018 Logistics - Monthly Meeting - Workshop - Internship - ASGSR Introduction of team leads and goals of each group Introduction of all members Q/A Three-tier Client Strategy to Democratize Data Physiological changes, pathway enrichment, differential expression, normalization, processing metadata, reproducibility, Data federation/integration with heterogeneous bioinformatics external databases The GLDS currently serves over 100 omics investigations to the biomedical community via open access. In order to expand the scope of metadata record searches via the GLDS, we designed a metadata warehouse that collects and updates metadata records from external systems housing similar data. To demonstrate the capabilities of federated search and retrieval of these data, we imported metadata records from three open-access data systems into the GLDS metadata warehouse: NCBI's Gene Expression Omnibus (GEO), EBI's PRoteomics IDEntifications (PRIDE) repository, and the Metagenomics Analysis server (MG-RAST). Each of these systems defines metadata for omics data sets differently. One solution to bridge such differences is to employ a common object model (COM) to which each systems' representation of metadata can be mapped. Warehoused metadata records are then transformed at ETL to this single, common representation. Queries generated via the GLDS are then executed against the warehouse, and matching records are shown in the COM representation (Fig. 1). While this approach is relatively straightforward to implement, the volume of the data in the omics domain presents challenges in dealing with latency and currency of records. Furthermore, the lack of a coordinated has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

GeneLab↗

Workplace Charging Data Collection and Behavior

This paper is a partner to the actual data being released. The data gathered and maintained by NREL tracks over 300 vehicles during the course of a 4-year period and how they behave in a workplace charging capacity. The data is further enriched by examining the effect of free charging versus paid charging. There is also a distinction in data marked by the onset of Covid-19. Vehicles are owned and operated by employees and range from smaller pack PHEV to larger pack BEVs.

33 ADVANCED PROPULSION SYSTEMS↗

Workplace Charging Data

A data set gathered and maintained by NREL that tracks over 300 vehicles during the course of a 4-year period and how they behave in a workplace charging capacity. The data is further enriched by examining the effect of free charging versus paid charging. There is also a distinction in data marked by the onset of Covid-19. Vehicles are owned and operated by employees and range from smaller pack PHEV to larger pack BEVs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Cosmic ray spectrum of protons plus helium nuclei between 6 and 158 TeV from HAWC data

Here, a measurement with high statistics of the differential energy spectrum of light elements in cosmic rays, in particular, of primary H plus He nuclei, is reported. The spectrum is presented in the energy range from 6 to 158 TeV per nucleus. Data was collected with the High Altitude Water Cherenkov (HAWC) Observatory between June 2015 and June 2019. The analysis was based on a Bayesian unfolding procedure, which was applied on a subsample of vertical HAWC data that was enriched to 82% of events induced by light nuclei. To achieve the mass separation, a cut on the lateral age of air shower data was set guided by predictions of CORSIKA/QGSJET-II-04 simulations. The measured spectrum is consistent with a broken power-law spectrum and shows a kneelike feature at around E = 24.0$^{+3.6}_{-3.1}$ TeV , with a spectral index γ = -2.51 ± 0.02 before the break and with γ = -2.83 ± 0.02 above it. The feature has a statistical significance of 4.1σ. Within systematic uncertainties, the significance of the spectral break is 0.8σ.

79 ASTRONOMY AND ASTROPHYSICS↗

Harnessing the Risk-Related Data Supply Chain: An Information Architecture Approach to Enriching Human System Research and Operations Knowledge

An Information Architecture facilitates the understanding and, hence, harnessing of the human system risk-related data supply chain which enhances the ability to securely collect, integrate, and share data assets that improve human system research and operations. By mapping the risk-related data flow from raw data to useable information and knowledge (think of it as a data supply chain), the Human Research Program (HRP) and Space Life Science Directorate (SLSD) are building an information architecture plan to leverage their existing, and often shared, IT infrastructure.

Buquo, Lynn E.↗

Harnessing the Risk-Related Data Supply Chain: An Information Architecture Approach to Enriching Human System Research and Operations Knowledge

NASA's Human Research Program (HRP) and Space Life Sciences Directorate (SLSD), not unlike many NASA organizations today, struggle with the inherent inefficiencies caused by dependencies on heterogeneous data systems and silos of data and information spread across decentralized discipline domains. The capture of operational and research-based data/information (both in-flight and ground-based) in disparate IT systems impedes the extent to which that data/information can be efficiently and securely shared, analyzed, and enriched into knowledge that directly and more rapidly supports HRP's research-focused human system risk mitigation efforts and SLSD s operationally oriented risk management efforts. As a result, an integrated effort is underway to more fully understand and document how specific sets of risk-related data/information are generated and used and in what IT systems that data/information currently resides. By mapping the risk-related data flow from raw data to useable information and knowledge (think of it as the data supply chain), HRP and SLSD are building an information architecture plan to leverage their existing, shared IT infrastructure. In addition, it is important to create a centralized structured tool to represent risks including attributes such as likelihood, consequence, contributing factors, and the evidence supporting the information in all these fields. Representing the risks in this way enables reasoning about the risks, e.g. revisiting a risk assessment when a mitigation strategy is unavailable, updating a risk assessment when new information becomes available, etc. Such a system also provides a concise way to communicate the risks both within the organization as well as with collaborators. Understanding and, hence, harnessing the human system risk-related data supply chain enhances both organizations' abilities to securely collect, integrate, and share data assets that improve human system research and operations.

Buquo, Lynn↗

Development of physics-consistent conditional diffusion model to overcome data scarcity in critical heat flux

Deep generative modeling provides a powerful pathway to overcome data scarcity in energy-related applications where experimental data are often limited. By learning the underlying probability distribution of the training dataset, deep generative models, such as the diffusion model, can generate high-fidelity synthetic samples that statistically resemble the training data. Such synthetic data generation can significantly enrich the size and diversity of the available training data, and more importantly, improve the robustness of downstream machine learning models in predictive tasks. The objective of this paper is to investigate the effectiveness of diffusion models for overcoming data scarcity in nuclear energy applications. By leveraging a public dataset on critical heat flux which covers a wide range of commercial nuclear reactor operational conditions, we developed a diffusion model that can generate an arbitrary amount of synthetic samples. Since a vanilla diffusion model can only generate samples randomly, we also developed a conditional diffusion model capable of generating targeted critical heat flux data under user-specified thermal-hydraulic conditions. The performance of the diffusion model was evaluated based on its ability to capture empirical feature distributions and pair-wise correlations, as well as to maintain physical consistency. The results showed that both the diffusion model and conditional diffusion model can successfully generate realistic and physics-consistent critical heat flux data. Furthermore, uncertainty quantification results demonstrate that the conditional diffusion model is highly effective in augmenting critical heat flux data while maintaining acceptable levels of uncertainty.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Gross and Net Soil Methane Flux and Ancillary Data, Edgewater, MD, USA, summer 2022

This data package contains measurements used to quantify methane cycling and environmental conditions in coastal forest soils during the 2022 growing season. It includes time‑series data of soil methane flux, soil respiration, soil temperature, and volumetric water content collected from soil monoliths transplanted along an inundation and salinity gradient. The package also provides one‑time measurements from a stable‑isotope pool‑dilution incubation, including gravimetric water content, methane headspace concentrations, and ¹³CH₄ enrichment over time. Data files are provided in comma‑separated values (CSV) format, with accompanying metadata and readme documentation in PDF and plain‑text formats. All files can be opened with standard software such as R, Python, or spreadsheet programs capable of handling CSV files. The metadata file describes variable definitions, units, processing steps, and the structure of each data table to support reuse and integration with other datasets.

13-C↗