Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data enrichment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Replication Data for: Quantifying oxygen induced surface enrichment of a dilute PdAu alloy catalyst

The data underlying this published work have been made publicly available in this repository as part of the IMASC Data Management Plan. This work was supported as part of the Integrated Mesoscale Architectures for Sustainable Catalysis (IMASC), an Energy Frontier Research Center funded by the U.S. Department of Energy, Office of Science, Basic Energy Sciences under Award # DE-SC0012573.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Eev (enrich Enforce Validate) With Cpefinder

This code is designed to take an existing STIX bundle with vulnerability data and enrich it with additional potential vulnerabilities to provide further insight during threat analysis. It also acts as a launch platform for other enrichments tools. The additional tools include, WAVgraph and STIXEnforcer.

Beckman, BryanR [Idaho National Laboratory (INL), ↗

Ontology-Enriched Specifications Enabling Findable, Accessible, Interoperable, and Reusable Marine Metagenomic Datasets in Cyberinfrastructure Systems

Marine microbial ecology requires the systematic comparison of biogeochemical and sequence data to analyze environmental influences on the distribution and variability of microbial communities. With ever-increasing quantities of metagenomic data, there is a growing need to make datasets Findable, Accessible, Interoperable, and Reusable (FAIR) across diverse ecosystems. FAIR data is essential to developing analytical frameworks that integrate microbiological, genomic, ecological, oceanographic, and computational methods. Although community standards defining the minimal metadata required to accompany sequence data exist, they haven’t been consistently used across projects, precluding interoperability. Moreover, these data are not machine-actionable or discoverable by cyberinfrastructure systems. By making ‘omic and physicochemical datasets FAIR to machine systems, we can enable sequence data discovery and reuse based on machine-readable descriptions of environments or physicochemical gradients. In this work, we developed a novel technical specification for dataset encapsulation for the FAIR reuse of marine metagenomic and physicochemical datasets within cyberinfrastructure systems. This includes using Frictionless Data Packages enriched with terminology from environmental and life-science ontologies to annotate measured variables, their units, and the measurement devices used. This approach was implemented in Planet Microbe, a cyberinfrastructure platform and marine metagenomic web-portal. Here, we discuss the data properties built into the specification to make global ocean datasets FAIR within the Planet Microbe portal. We additionally discuss the selection of, and contributions to marine-science ontologies used within the specification. Finally, we use the system to discover data by which to answer various biological questions about environments, physicochemical gradients, and microbial communities in meta-analyses. This work represents a future direction in marine metagenomic research by proposing a specification for FAIR dataset encapsulation that, if adopted within cyberinfrastructure systems, would automate the discovery, exchange, and re-use of data needed to answer broader reaching questions than originally intended.

59 BASIC BIOLOGICAL SCIENCES↗

Search for CP violation in t$\overline{\textrm{t}}$H and tH production in multilepton channels in proton-proton collisions at $\sqrt{s}$ = 13 TeV

The charge-parity (CP) structure of the Yukawa interaction between the Higgs (H) boson and the top quark is measured in a data sample enriched in the t$\overline{t}$H and tH associated production, using 138 fb -1 of data collected in proton-proton collisions at $\sqrt{s}$ = 13 TeV by the CMS experiment at the CERN LHC. The study targets events where the H boson decays via H → WW or H → ττ and the top quarks decay via t → Wb: the W bosons decay either leptonically or hadronically, and final states characterized by the presence of at least two leptons are studied. Machine learning techniques are applied to these final states to enhance the separation of CP -even from CP -odd scenarios. Two-dimensional confidence regions are set on $κ$ t and $\widetilde{k}$t, which are respectively defined as the CP -even and CP -odd top-Higgs Yukawa coupling modifiers. No significant fractional CP -odd contributions, parameterized by the quantity |$f^{Htt}_{CP}$| are observed; the parameter is determined to be |$f^{Htt}_{CP}$| = 0.59 with an interval of (0.24, 0.81) at 68% confidence level. The results are combined with previous results covering the H → ZZ and H → γγ decay modes, yielding two- and one-dimensional confidence regions on $κ$ t and $\widetilde{k}$t, while |$f^{Htt}_{CP}$| is determined to be |$f^{Htt}_{CP}$| = 0.28 with an interval of |$f^{Htt}_{CP}$| < 0.55 at 68% confidence level, in agreement with the standard model CP -even prediction of |$f^{Htt}_{CP}$| = 0.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Measurement of the top quark mass using events with a single reconstructed top quark in pp collisions at $$ \sqrt{s} $$ = 13 TeV

A measurement of the top quark mass is performed using a data sample enriched with single top quark events produced in the t channel. The study is based on proton- proton collision data, corresponding to an integrated luminosity of 35.9 fb -1 , recorded at √s = 13 TeV by the CMS experiment at the LHC in 2016. Candidate events are selected by requiring an isolated high-momentum lepton (muon or electron) and exactly two jets, of which one is identified as originating from a bottom quark. Multivariate discriminants are designed to separate the signal from the background. Optimized thresholds are placed on the discriminant outputs to obtain an event sample with high signal purity. The top quark mass is found to be $172.13^{+0.76}_{-0.77}$ GeV, where the uncertainty includes both the statistical and systematic components, reaching sub-GeV precision for the first time in this event topology. The masses of the top quark and antiquark are also determined separately using the lepton charge in the final state, from which the mass ratio and difference are determined to be $0.9952^{+0.0079}_{-0.0104}$ and $0.83^{+1.79}_{-1.35}$ GeV, respectively. The results are consistent with CPT invariance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SeismoGen: Seismic Waveform Synthesis Using GAN With Application to Seismic Data Augmentation

Abstract Detecting earthquake arrivals within seismic time series can be a challenging task. Visual, human detection has long been considered the gold standard but requires intensive manual labor that scales poorly to large data sets. In recent years, automatic detection methods based on machine learning have been developed to improve the accuracy and efficiency. However, the accuracy of those methods relies on access to a sufficient amount of high‐quality labeled training data, often tens of thousands of records or more. We aim to resolve this dilemma by answering two questions: (1) provided with a limited amount of reliable labeled data, can we use them to generate additional, realistic synthetic waveform data? and (2) can we use those synthetic data to further enrich the training set through data augmentation, thereby enhancing detection algorithms? To address these questions, we use a generative adversarial network (GAN), a type of machine learning model which has shown supreme capability in generating high‐quality synthetic samples in multiple domains. Once trained, our GAN model is capable of producing realistic seismic waveforms of multiple labels (noise and event classes). Applied to real Earth seismic data sets in Oklahoma, we show that data augmentation from our GAN‐generated synthetic waveforms can be used to improve earthquake detection algorithms in instances when only small amounts of labeled training data are available.

Wang, Tiantong↗

3D Continuous Forcing Dataset from 3D Constrained Variational Analysis at SGP

The continuous 3D large-scale forcing (VARANAL3D) data set derived from 3D constrained variational analysis (3DCVA) extends the conventional constrained variational analysis method by incorporating multiple sub-columns within the analysis domain. This advancement introduces spatial variability into the large-scale forcing fields, thereby enriching the data set’s applicability. The VARANAL3D data set spans from 2004 to 2018 and covers a region of 5˚×4.5˚ domain around the ARM SGP site. The analysis domain is divided into 10×9 sub-columns with 0.5˚ resolution. The 3D large-scale forcing data provides necessary variables to drive and evaluate single-column models (SCM), cloud-resolving models (CRM) ,and large-eddy simulations (LES), as well as information for testing model sensitivity to spatial variability of the large-scale forcing data, facilitating more rigorous testing and refinement of physical processes in SCM/CRM/LES.

54 ENVIRONMENTAL SCIENCES↗

Search for pair production of heavy particles decaying to a top quark and a gluon in the lepton+jets final state in proton–proton collisions at $\sqrt{s}=13\,\text {Te}\hspace{-.08em}\text {V}$

A search is presented for the pair production of new heavy resonances, each decaying into a top quark (t) or antiquark and a gluon (g). The analysis uses data recorded with the CMS detector from proton–proton collisions at a center-of-mass energy of 13 TeV at the LHC, corresponding to an integrated luminosity of 138 fb -1 . Events with one muon or electron, multiple jets, and missing transverse momentum are selected. After using a deep neural network to enrich the data sample with signal-like events, distributions in the scalar sum of the transverse momenta of all reconstructed objects are analyzed in the search for a signal. No significant deviations from the standard model prediction are found. Upper limits at 95% confidence level are set on the product of cross section and branching fraction squared for the pair production of excited top quarks in the t* → tg decay channel. The upper limits range from 120 to 0.8 fb for a t* with spin-1/2 and from 15 to 1.0 fb for a t* with spin-3/2. These correspond to mass exclusion limits up to 1050 and 1700 GeV for spin-1/2 and spin-3/2 t* particles, respectively. These are the most stringent limits to date on the existence of t* → tg resonances.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A combinatorially complete epistatic fitness landscape in an enzyme active site

Protein engineering often targets binding pockets or active sites which are enriched in epistasis—nonadditive interactions between amino acid substitutions—and where the combined effects of multiple single substitutions are difficult to predict. Few existing sequence-fitness datasets capture epistasis at large scale, especially for enzyme catalysis, limiting the development and assessment of model-guided enzyme engineering approaches. We present here a combinatorially complete, 160,000-variant fitness landscape across four residues in the active site of an enzyme. Assaying the native reaction of a thermostable β-subunit of tryptophan synthase (TrpB) in a nonnative environment yielded a landscape characterized by significant epistasis and many local optima. These effects prevent simulated directed evolution approaches from efficiently reaching the global optimum. There is nonetheless wide variability in the effectiveness of different directed evolution approaches, which together provide experimental benchmarks for computational and machine learning workflows. The most-fit TrpB variants contain a substitution that is nearly absent in natural TrpB sequences—a result that conservation-based predictions would not capture. Thus, although fitness prediction using evolutionary data can enrich in more-active variants, these approaches struggle to identify and differentiate among the most-active variants, even for this near-native function. Overall, this work presents a large-scale testing ground for model-guided enzyme engineering and suggests that efficient navigation of epistatic fitness landscapes can be improved by advances in both machine learning and physical modeling.

biocatalysis↗

A multi-ancestry genetic study of pain intensity in 598,339 veterans

Chronic pain is a common problem, with more than one-fifth of adult Americans reporting pain daily or on most days. It adversely affects the quality of life and imposes substantial personal and economic costs. Efforts to treat chronic pain using opioids had a central role in precipitating the opioid crisis. Despite an estimated heritability of 25–50%, the genetic architecture of chronic pain is not well-characterized, in part because studies have largely been limited to samples of European ancestry. To help address this knowledge gap, we conducted a cross-ancestry meta-analysis of pain intensity in 598,339 participants in the Million Veteran Program, which identified 126 independent genetic loci, 69 of which are new. Pain intensity was genetically correlated with other pain phenotypes, level of substance use and substance use disorders, other psychiatric traits, education level and cognitive traits. Integration of the genome-wide association studies findings with functional genomics data shows enrichment for putatively causal genes (n = 142) and proteins (n = 14) expressed in brain tissues, specifically in GABAergic neurons. Drug repurposing analysis identified anticonvulsants, β-blockers and calcium-channel blockers, among other drug groups, as having potential analgesic effects. Our results provide insights into key molecular contributors to the experience of pain and highlight attractive drug targets.

59 BASIC BIOLOGICAL SCIENCES↗

Calibration of a soft secondary vertex tagger using proton-proton collisions at s = 13 TeV with the ATLAS detector

Several processes studied by the ATLAS experiment at the Large Hadron Collider produce low-momentum b-flavored hadrons in the final state. This paper describes the calibration of a dedicated tagging algorithm that identifies b-flavored hadrons outside of hadronic jets by reconstructing the soft secondary vertices originating from their decays. The calibration is based on a proton-proton collision dataset at a center-of-mass energy of 13 TeV corresponding to an integrated luminosity of 140 fb -1 . Scale factors used to correct the algorithm’s performance in simulated events are extracted for the b-tagging efficiency and the mistag rate of the algorithm using a data sample enriched in $t\overline{t}$ events. Several orthogonal measurement regions are defined, binned as a function of the multiplicities of soft secondary vertices and jets containing a b-flavored hadron in the event. The mistag rate scale factors are estimated separately for events with low and high average numbers of interactions per bunch crossing. The results, which are derived from events with low missing transverse momentum, are successfully validated in a phase space characterized by high missing transverse momentum and therefore are applicable to new physics searches carried out in either phase space regime.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Using DNA affinity purification sequencing (DAP-seq) to identify in vitro binding sites of potential Novosphingobium aromaticivorans DSM12444 transcription factors

Genome-wide binding sites of 44 putative transcription factors (TFs) from Novosphingobium aromaticivorans DSM12444 were analyzed using DNA affinity purification sequencing. We report that 32 of these TFs have at least one area of enrichment. These data will help better understand aromatic metabolism and other features of N. aromaticivorans biology.

DAP-seq↗

Using DNA Affinity Purification sequencing (DAP-seq) to identify in vitro binding sites of transcription factors potentially involved in aromatic degradation

The genome-wide binding sites of 44 transcription factors from the aromatic metabolizing Alphaproteobacterium Novosphingobium aromaticivorans were identified using DNA Affinity Purification sequencing (DAP-seq). We report 32 of these transcription factors have at least one area of enrichment. These data will be valuable for better understanding of aromatic metabolism.

aromatic metabolism↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗

Genome-resolved metagenomics reveals role of iron metabolism in drought-induced rhizosphere microbiome dynamics

Recent studies have demonstrated that drought leads to dramatic, highly conserved shifts in the root microbiome. At present, the molecular mechanisms underlying these responses remain largely uncharacterized. Here we employ genome-resolved metagenomics and comparative genomics to demonstrate that carbohydrate and secondary metabolite transport functionalities are overrepresented within drought-enriched taxa. These data also reveal that bacterial iron transport and metabolism functionality is highly correlated with drought enrichment. Using time-series root RNA-Seq data, we demonstrate that iron homeostasis within the root is impacted by drought stress, and that loss of a plant phytosiderophore iron transporter impacts microbial community composition, leading to significant increases in the drought-enriched lineage, Actinobacteria. Finally, we show that exogenous application of iron disrupts the drought-induced enrichment of Actinobacteria, as well as their improvement in host phenotype during drought stress. Collectively, our findings implicate iron metabolism in the root microbiome’s response to drought and may inform efforts to improve plant drought tolerance to increase food security.

59 BASIC BIOLOGICAL SCIENCES↗

Two decades of fumigation data from the Soybean Free Air Concentration Enrichment facility

Abstract The Soybean Free Air Concentration Enrichment (SoyFACE) facility is the longest running open-air carbon dioxide and ozone enrichment facility in the world. For over two decades, soybean, maize, and other crops have been exposed to the elevated carbon dioxide and ozone concentrations anticipated for late this century. The facility, located in East Central Illinois, USA, exposes crops to different atmospheric concentrations in replicated octagonal ~280 m 2 Free Air Concentration Enrichment (FACE) treatment plots. Each FACE plot is paired with an untreated control (ambient) plot. The experiment provides important ground truth data for predicting future crop productivity. Fumigation data from SoyFACE were collected every four seconds throughout each growing season for over two decades. Here, we organize, quality control, and collate 20 years of data to facilitate trend analysis and crop modeling efforts. This paper provides the rationale for and a description of the SoyFACE experiments, along with a summary of the fumigation data and collation process, weather and ambient data collection procedures, and explanations of air pollution metrics and calculations.

60 APPLIED LIFE SCIENCES↗

R -Matrix Analysis and Statistical Properties of Dysprosium Isotopes in the Neutron Energy Ranges Up To A Few Kev

In support of the Nuclear Criticality Safety Program, a set of evaluated resonance parameters was generated for seven dysprosium isotopes in the neutron energy range from thermal up to a few keV. The evaluation methodology used the Reich-Moore approximation to fit, with the R-matrix code SAMMY, the high-resolution capture and transmission measurements on natural and enriched samples recently performed at the Rensselaer Polytechnic Institute Gaerttner LINear ACcelerator facility. Additional transmission data measured on enriched samples by Liou at the Columbia University Nevis synchrocyclotron in the mid-seventies were used to gauge the neutron widths above 15 eV. Thermal constants such as absorption and (in)coherent scattering cross sections and corresponding scattering lengths were calibrated to the National Institute of Standards and Technology’s compilation except for 161,164 Dy isotopes.

07 ISOTOPE AND RADIATION SOURCES↗

A Review of Candidates for a Validation Data Set for High-Assay Low-Enrichment Uranium Fuels

Many advanced reactor concept designs rely on high-assay low-enriched uranium (HALEU) fuel, enriched up to approximately 19.75% 235 U by weight. Efforts are underway by the US government to increase HALEU production in the United States to meet anticipated needs. However, very few data exist for validation of computational models that include HALEU, beyond a few fresh fuel benchmark specifications in the International Reactor Physics Experiment Evaluation Project. Nevertheless, there are other data with potential value available for developing into quality benchmarks for use in data- and software-validation efforts. This paper reviews the available evaluated HALEU fuel benchmarks and some of the potentially relevant benchmarks for fresh highly enriched uranium. It then introduces experimental data for HALEU fuel irradiated at Idaho National Laboratory, from relatively recent irradiation programs at the Advanced Test Reactor. Such data should be evaluated and, if valuable, collected into detailed benchmark specifications to meet the needs of HALEU-based reactor designers.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗