Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data discover”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

EDXplorer: A Utility for APA Analysis

The automated particle analysis (APA) method of scanning electron microscopy (SEM) energy dispersive X-ray spectroscopy (EDS/EDX) is a useful tool for analyzing the elemental and morphological data of particulate samples. Often, such datasets have many thousands of particles, and it can be difficult to sift through the data to find meaningful trends. EDXplorer is a software utility for processing APA data. They enable the user to easily load data and determine the important components and aspects of the datasets using a powerful and versatile library of data plotting functions, mainly centered around scatter plots and histograms. These programs are intended to fill a void in the data processing of data from certain instruments, where often the user must rely on their own code or other software that is not user-friendly. EDXplorer is intended as a general plotting utility for browsing through data and discovering data trends.

Moseley, Duncan [ORNL] (ORCID:0000000343518347)

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives.

access

Empowering Geothermal Research: The Geothermal Data Repository's New AI Research Assistant

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has integrated a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets to create an Artificially Intelligent (AI) research assistant. By leveraging work done to make GDR metadata machine-readable and an open-source LLM integration model called the Energy Language Model, developed by the National Renewable Energy Laboratory, AskGDR serves as a virtual research assistant to GDR users. It provides answers to a variety of user-provided questions using natural language processing and generative machine learning. Users can get answers to questions about specific datasets, including inquiries about the equipment, assumptions and methodologies used in the origination of the data; or more abstract questions, such as the applicability of data to specific research fields. AskGDR improves the discoverability of geothermal data by helping guide users to datasets beyond simple keyword searches. It enables users to find data based on properties of the data, discover information contained within supporting documents, and explore data from projects related to their research objectives. This paper will outline the development, integration, output, and efficacy of the AskGDR LLM, including adherence to scientific rigor through improvements designed to increase the accuracy of generated answers, avoid speculation, and provide proper references for all resources used.

access

Semi-Analytical Hierarchical Bayesian Inference of Nonlinear Model Structure in Stochastic Dynamics: Applied to Compartmental Models of Infectious Diseases

A Bayesian computational framework for parsimonious inference in stochastic nonlinear dynamical systems is presented. This framework enables the concurrent estimation of system states, time-varying parameters, time-invariant parameters, and the optimal sparsity structure of the model parameters. Because differential equation-based models are often simplified mechanistic or phenomenological representations, robust inference from noisy measurement data requires explicit treatment of model error and uncertainty. Model error and time-varying parameters can be represented as random processes, enabling inference while making minimal assumptions about the underlying sources of discrepancy and variability. Adopting stochastic differential equation representations affords the model significant flexibility, but can also render it susceptible to overfitting during statistical inversion, where the inferred model may track noise rather than the underlying signal. To alleviate the effects of overfitting and to enable the discovery of the optimal sparse representation of the time-invariant parameters, a Bayesian sparse learning algorithm is embedded within the framework. This sparse learning framework adopts an approximate hierarchical Bayesian setting defined by a series of semi-analytical expressions. The model structure inference framework is validated using a stochastic compartmental model for tracking and forecasting active cases of an infectious disease. Compartmental models describe population-level infectious disease dynamics through interactions among population fractions grouped by disease state. Mathematically, such models consist of a system of coupled ordinary differential equations. This example adopts an expressive compartmental model that includes multiple possible interactions between disease states, motivated by early uncertainty surrounding COVID-19 reinfection dynamics and their implications for long-term epidemic forecasting. The sparse learning exercise permits the inference of a priori unknown epidemiological dynamics from simulated public health data, discovering the nested compartmental model that optimizes the trade-off between average data-fit and model complexity. It is shown that inducing sparsity among the model parameters eliminates redundant interactions between compartments, equivalently revealing the optimal coupling structure between differential equations.

97 MATHEMATICS AND COMPUTING

Multimetallic Layered Composites (MMLCs) for Rapid, Economical Advanced Reactor Deployment (Final Report)

This project focused on the development of multi-metallic layered composites (MMLCs) for advanced fission reactor technologies. There are many instances where one alloy or material simply cannot meet all the demands thrown at it by a reactor system, or cannot allow it to perform as strongly as one would like. Instead of focusing all our effort on developing one perfect alloy, we seek to leverage the design principle of “separation of functionality,” used in many other arenas in design, to boost performance beyond single alloys alone. One illustrative example shows the power of this approach for molten salt-cooled reactors: A three meter tall, three meter diameter reactor vessel made of Incoloy 800 was quoted at $\$$500k in 2018. A Hastelloy N vessel was quoted at $\$$5M. An MMLC vessel, in which a layer of Hastelloy N would be weld-overlaid onto Incoloy 800, was quoted at $\$$700k, and it would achieve the same performance. The potential economic gains of leveraging this approach are therefore substantial. At a minimum, each MMLC would contain one core structural layer and one coolant-facing corrosion-resistant layer. Sometimes, MMLCs required buffer layers, as the structural and corrosion-resistant layers were metallurgically incompatible. In other words, they didn’t always play nice, thus separating layers compatible with both functioned as intermediaries to keep the composite together. However, in doing so we inevitably produce new interfaces, where new issues can arise. Therefore, this project focused on what happens at these interfaces from a combination of high temperatures, irradiation, corrosion, and time. After all, a reactor makes money when it is operating, and outages of any kind erode its economic viability. First, we set out to experimentally prove that MMLCs for at least two advanced reactor systems can be made, today, in US domestic facilities. In this respect we were successful – one MMLC (a Ni-201/Incoloy 800H composite) was successfully made and drawn into two-inch coolant piping. Others were attempted, though new issues relating to cracking in vanadium layers for one and radiation damage performance of the corrosion-resistant layer in another prevented us from moving further in those specific arenas – these are engineering problems which deserve continued focus after this project. Additional experimental work focused on long-term corrosion testing of the outermost layers of the salt-cooled and liquid lead-cooled MMLC concepts, which would then be fed into predictions of how long the MMLCs could last. Next, computational (thermodynamics and atomistic) simulation studies studied how much we expect the interfaces to “blend,” due to the mixing action of neutron irradiation. This eats into both the margin for the structural layer of each MMLC, as dilution from the corrosion-resistant layer into the structural layer would decrease the total load-bearing capacity of an MMLC of finite size. On the other hand, dilution of the corrosion-resistant layer into the structural layer further reduced the margin of corrodible material, reducing the lifetime of the MMLC or necessitating extra thickness to be imparted to the MMLC to meet its functional requirements. Work here focused on irradiation-induced segregation to predict new phases which may embrittle the MMLCs, as well as quantifying irradiation-induced mixing at each interface. The results showed that mixing is expected, but it is both steady and therefore predictable, and not lifetime-limiting for most MMLC concepts – it simply has to be accounted for in calculations of reactor performance when utilizing an MMLC. Then, full-core simulations using the experimentally-derived corrosion data, the computationally discovered irradiation-induced mixing data (partially validated by experiment), and existing, benchmarked core designs for large and small sized reactor concepts (one salt-cooled, one lead-cooled) were conducted to quantify any expansion of reactor operating envelopes achieved by utilizing these MMLCs. This new framework, called REX (Reactor Envelope Expansion), incorporates a combination of core neutronics, thermal hydraulics, and the material performance data derived from this project to see how using an MMLC expands advanced fission reactor operating envelopes. It was discovered that in some cases, MMLC utilization does indeed increase the maximum operating temperatures and cycle lengths of reactor concepts, while in other cases it does not. Finally, our tech-to-market (T2M) strategy was not necessarily to create specific embodiments of MMLCs for immediate sale (because getting into the nuclear market is incredibly slow and laden with regulation, this is a long-term goal), but rather immediate stimulation of US industry using the design approach of MMLCs derived from this project. In this respect we were successful, as one of the PhD students funded on this project co-founded Allium Engineering, Inc., which created a stainless steel / low-alloy steel MMLC to function as chloride corrosion-resistant rebar for embedding into concrete structures. Allium Engineering continues to be successful, having recently opened their first factory as of this writing.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Structural and functional analyses of SARS-CoV-2 Nsp3 and its specific interactions with the 5’ UTR of the viral genome

ABSTRACT Non-structural protein 3 (Nsp3) is the largest open reading frame encoded in the SARS-CoV-2 genome, essential for the formation of double-membrane vesicles (DMV) wherein viral RNA replication occurs. We conducted an extensive structure-function analysis of Nsp3 and determined the crystal structures of the ubiquitin-like 1 (Ubl1), nucleic acid binding (NAB), β-coronavirus-specific marker (βSM) domains, and a sub-region of the Y domain of this protein. We show that the Ubl1, ADP-ribose phosphatase (ADRP), human SARS Unique (HSUD), NAB, and Y domains of Nsp3 bind the 5’ UTR of the viral genome and that the Ubl1 and Y domains possess affinity for recognition of this region, suggesting high specificity. The Ubl1-Nucleocapsid (N) protein complex binds the 5’ UTR with greater affinity than the individual proteins alone. Our results suggest that multiple domains of Nsp3, particularly Ubl1 and Y, shepherd the 5’ UTR of the viral genome during translocation through the DMV membrane, priming the Ubl1 domain to load the genome onto N protein. IMPORTANCE The largest protein encoded by the SARS-CoV-2 genome is Nsp3. In infected cells, this multi-domain protein forms a pore structure in the virus-induced double-membrane vesicles (DMV). We have incomplete data on Nsp3 molecular structure, and here, we describe crystal structures for multiple domains of Nsp3. It is thought that newly replicated viral RNA transits through the DMV pore; however, we possess incomplete data on which regions of Nsp3 actually interact with RNA. Here, we present data showing that five domains of Nsp3 interact with the 5’ UTR of the SARS-CoV-2 RNA, including the Y domain for which no function has ever been discovered. These data suggest that the pore structure plays an active role in recognizing the terminal end of the genome, transiting and loading the viral RNA onto the cytoplasmic nucleocapsid protein. These data help expand our knowledge of Nsp3 structure and function and the SARS-CoV-2 replication cycle.

Microbiology

Shapiro steps observed in a two-dimensional Yukawa solid modulated by a one-dimensional vibrational periodic substrate

Depinning dynamics of a two-dimensional (2D) solid dusty plasma modulated by a one-dimensional (1D) vibrational periodic substrate are investigated using Langevin dynamical simulations. As the uniform driving force increases gradually, from the overall drift velocity varying with the driving force, four significant Shapiro steps are discovered. The data analysis results indicate that, when the ratio of the frequency from the drift motion over potential wells to the external frequency from the modulation substrate is close to integers, dynamic mode locking occurs, corresponding to the discovered Shapiro steps. In conclusion, around both termini of the first and fourth Shapiro steps, the transitions are found to be always continuous, however, the transition between the second and third Shapiro steps is discontinuous, probably due to the different arrangements of particles.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

BrainXcan identifies brain features associated with behavioral and psychiatric traits using large-scale genetic and imaging data

Advances in brain MRI have enabled many discoveries in neuroscience. Case-control comparisons of brain MRI features have highlighted potential causes of psychiatric and behavioral disorders. However, due to the cost and difficulty of collecting MRI data, most studies have small sample sizes, limiting their reliability. Furthermore, reverse causality complicates interpretation because many observed brain differences are the result rather than the cause of the disease. Here we propose a method (BrainXcan) that leverages the power of large-scale genomewide association studies (GWAS) and reference brain MRI data to discover new mechanisms of disease etiology and validate existing ones. BrainXcan tests the association with genetic predictors of brain MRI-derived features and complex traits to pinpoint relevant brain-wide and region-specific features. Requiring only genetic data, BrainXcan allows us to test a host of hypotheses on mental illness, across many MRI modalities, using public data resources. For example, our method shows that reduced axonal density across the brain is associated with schizophrenia risk, consistent with the disconnectivity hypothesis. We also find that the hippocampus volume is associated with schizophrenia risk, highlighting the potential of our approach. Taken together, our results show the promise of BrainXcan to provide insights into the biology of GWAS traits.

Association study

Atmospheric wind energization of ocean weather

Ocean weather comprises vortical and straining mesoscale motions, which play fundamentally different roles in the ocean circulation and climate system. Vorticity determines the movement of major ocean currents and gyres. Strain contributes to frontogenesis and the deformation of water masses, driving much of the mixing and vertical transport in the upper ocean. While recent studies have shown that interactions with the atmosphere damp the ocean’s mesoscale vortices O(100) km in size, the effect of winds on straining motions remains unexplored. Here, we derive a theory for wind work on the ocean’s vorticity and strain. Using satellite and model data, we discover that wind damps strain and vorticity at an equal rate globally, and unveil striking asymmetries based on their polarity. Subtropical winds damp oceanic cyclones and energize anticyclones outside strong current regions, while subpolar winds have the opposite effect. A similar pattern emerges for oceanic strain, where subtropical convergent flow is damped along the west-equatorward east-poleward direction and energized along the east-equatorward west-poleward direction. These findings reveal energy pathways through which the atmosphere shapes ocean weather.

54 ENVIRONMENTAL SCIENCES

Predicting the viscoplastic response of a crystallizing fluoropolymer using transient network theory

We employ a molecular theory of dynamic polymer networks to describe the viscoplastic response of rubbery FK-800, a thermoplastic copolymer of chlorotrifluoroethylene and vinylidene fluoride, over a broad range of thermal histories. The kinetics of crystallization at different annealing temperatures was modeled using a modified Avrami equation, whose parameters were found to evolve through simple relationships over the full temperature range of the rubbery state. By fitting experimental compression data, we discovered predictable trends for the physical parameters in our mechanical model over its full range of crystallinities (up to ≈20%) and provided insights based on molecular-level physics to justify them. Using this, an end-to-end model was developed to predict the yielding and post-yield behavior of rubbery FK-800 for arbitrary thermal histories. The model successfully predicted the highly nonlinear evolution of characteristic mechanical signatures (stiffness, yield point, post-yield drop) throughout the crystallization process. A statistical analysis of variance test was employed to determine that the measured variations in the mechanical behavior of rubbery FK-800 are primarily dictated by its fractional crystallinity, regardless of its exact thermal history.

36 MATERIALS SCIENCE

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES

Final technical report for DE-SC0022255: Discovering Physically Meaningful Structures from Climate Extreme Data

The past two decades have witnessed natural disasters and extreme weather events that affect millions of people. At the same time, the data volume from high-resolution climate models, satellite, in-situ and ground-based measurements have substantially increased to petabyte scales. These new and readily accessible datasets create the previously missing pipeline required for scientific machine learning (ML) and therefore new opportunities for improved understanding and prediction capability of climate extreme events. This project developed a deep latent variable model framework to discover physically meaningful hidden structures from high-dimensional, spatiotemporal climate extreme data.

97 MATHEMATICS AND COMPUTING

Discovering methylated DNA motifs in bacterial nanopore sequencing data with MIJAMP

Abstract Bacterial DNA methylation is involved in diverse cellular functions, including modulation of gene expression, DNA repair, and restriction–modification systems for defense against viruses and other foreign DNA. Restriction systems hinder efforts to engineer organisms to produce fuels and chemicals from waste and renewable feedstocks by degrading DNA during transformation. Methylome analysis allows identification of motifs within a bacterial chromosome that may be targeted by native restriction enzymes. Further expression of the corresponding methyltransferases in Escherichia coli allows plasmid DNA to be protected from restriction in the target organism, thereby drastically enhancing transformation efficiency. Nanopore sequencing can detect methylated bases, but software is needed to transform modified base coordinates into methylated motifs. Here, we develop MIJAMP (MIJAMP Is Just A MethylBED Parser), a software package that was developed to discover methylated motifs from the output of ONT’s Modkit or other data in the methylBED format. MIJAMP employs a human-driven refinement strategy that empirically validates all motifs against genome-wide methylation data, thus eliminating incorrect motifs. MIJAMP also reports methylation data on specific, user-defined motifs. Using MIJAMP, we determined the methylated motifs both in a control strain (wild-type E. coli) and in Synecococcus sp. strain PCC7002, laying the foundation for improved transformation in this organism. MIJAMP is available at https://code.ornl.gov/alexander-public/mijamp/. One Sentence Summary: Here we describe software written to discover DNA methylation motifs from nanopore sequencing data.

59 BASIC BIOLOGICAL SCIENCES

Discovering Physically Meaningful Structures from Climate Extreme Data

The original proposal described an interdisciplinary team spanning UC San Diego (lead), Columbia University, and UC Irvine, with Columbia investigators including Pierre Gentine, Elias Bareinboim, and Marcus van Lier-Walqui. The proposal further specified a leadership structure in which Columbia co-investigators contributed across the three aims, with Co-PI Gentine serving as a point of contact with science teams and with responsibilities distributed across aims.

42 ENGINEERING

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL),

DELVE-ing into the Milky Way’s Globular Clusters: Assessing Extratidal Features in NGC 5897, NGC 7492, and Testing Detectability with Deeper Photometry

Extratidal features around globular clusters (GCs) are tracers of their disruption, stellar stream formation, and their host’s gravitational potential. However, these features remain challenging to detect due to their low surface brightness. We conduct a systematic search for such features around 19 GCs in the DECam Local Volume Exploration (DELVE) survey Data Release 2, discovering a new extra-tidal envelope around NGC 5897 and find tentative evidence for an extended envelope surrounding NGC 7492. Through a combination of dynamical modeling and analyzing synthetic stellar populations, we demonstrate these envelopes may have formed through tidal disruption. We use these models to explore the detectability of these features in the upcoming Legacy Survey of Space and Time (LSST), finding that while LSST’s deeper photometry will enhance detection significance, additional methods for foreground removal like proper motions or metallicities may be important for robust stream detection. Our results both add to the sample of globular clusters with extratidal features and provide insights on interpreting similar features in current and upcoming data.

Chiti, A. [Univ. of Chicago, IL (United States); S

Bioactivity Profiling of Chemical Mixtures for Hazard Characterization

Abstract The assessment and regulation of chemical toxicity to protect human health and the environment are done one chemical at a time and seldom at environmentally relevant concentrations. However, chemicals are found in the environment as mixtures, and their toxicity is largely unknown. Understanding the hazard posed by chemicals within the mixture is critical to enforce protective measures. Here, we demonstrate the application of bioactivity profiling of environmental water samples using the sentinel and ecotoxicology model species Daphnia to reveal the biomolecular response induced by exposure to real-world mixtures. We exposed a Daphnia strain to 30 sampled waters of the Chaobai River and measured the gene expression response profiles. Using a multiblock correlation analysis, we establish correlations between chemical mixtures identified in 30 water samples with gene expression patterns induced by these chemical mixtures. We identified 80 metabolic pathways putatively activated by mixtures of inorganic ions, heavy metals, polycyclic aromatic hydrocarbons, industrial chemicals, and a set of biocides, pesticides, and pharmacologically active substances. Our data-driven approach discovered both known bioactivity signatures with previously described modes of action and new pathways linked to undiscovered potential hazards. This study demonstrates the feasibility of reducing the complexity of real-world mixture toxicity to characterize the biomolecular effects of a defined number of chemical components based on gene expression monitoring of the sentinel species Daphnia.

Engineering

Data for reproducing the figures of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas

This deposit contains the raw data for reproducing research results of the paper Multimodal Super-Resolution: Discovering hidden physics and its application to fusion plasmas. The main contribution of this work is to utilize machine learning techniques to reconstruct and enhance the resolution of a diagnostic measurement from other available diagnostics in a system. The proposed techniques is called Diag2Diag.

diag2diag