Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Knowledge Graph”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Managing a project's legacy: implications for organizations and project management

Organizations that rely on projects to implement their products must find effective mechanisms for propagating lessons learned on one project throughout the organization. A broad view of what constitutes a project's 'legacy' is presented that includes not just the design products and leftover parts, but new processes, relationships, technology, skills, planning data, and performance metrics. Based on research evaluating knowledge reuse in innovative contexts, this paper presents an approach to project legacy management that focuses on collecting and using legacy knowledge to promote organizational learning and effective reuse, while addressing factors of post-project responsibility, information obsolescence, and the importance of ancillary contextual information. .

risk management↗

Semi-Automated, Object-Based Tomography of Dislocation Structures

The characterization of the three-dimensional arrangement of dislocations is important for many analyses in materials science. Dislocation tomography in transmission electron microscopy is conventionally accomplished through intensity-based reconstruction algorithms. Although such methods work successfully, a disadvantage is that they require many images to be collected over a large tilt range. Here, we present an alternative, semi-automated object-based approach that reduces the data collection requirements by drawing on the prior knowledge that dislocations are line objects. Our approach consists of three steps: (1) initial extraction of dislocation line objects from the individual frames, (2) alignment and matching of these objects across the frames in the tilt series, and (3) tomographic reconstruction to determine the full three-dimensional configuration of the dislocations. Drawing on innovations in graph theory, we employ a node-line segment representation for the dislocation lines and a novel arc-length mapping scheme to relate the dislocations to each other across the images in the tilt series. We demonstrate the method for a dataset collected from a dislocation network imaged by diffraction-contrast scanning transmission electron microscopy. Based on these results and a detailed uncertainty analysis for the algorithm, we discuss opportunities for optimizing data collection and further automating the method.

47 OTHER INSTRUMENTATION↗

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference↗

ARENA: Asynchronous Reconfigurable Accelerator Ring to Enable Data-Centric Parallel Computing

The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications. In this work, we propose ARENA – an asynchronous reconfigurable accelerator ring architecture as a potential scenario on how the future HPC and data centers will be like. Despite using the coarse-grained reconfigurable arrays (CGRAs) as the substrate platform, our key contribution is not only the CGRA-cluster design itself, but also the ensemble of a new architecture and programming model that enables asynchronous tasking across a cluster of reconfigurable nodes, so as to bring specialized computation to the data rather than the reverse. We presume distributed data storage without asserting any prior knowledge on the data distribution. Hardware specialization occurs at runtime when a task finds the majority of data it requires are available at the present node. In other words, we dynamically generate specialized CGRA accelerators where the data reside. The asynchronous tasking for bringing computation to data is achieved by circulating the task token, which describes the dataflow graphs to be executed for a task, among the CGRA cluster connected by a fast ring network. Evaluations on a set of HPC and data-driven applications across different domains show that ARENA can provide better parallel scalability with reduced data movement (53.9 percent). Compared with contemporary compute-centric parallel models, ARENA can bring on average 4.37× speedup. The synthesized CGRAs and their task-dispatchers only occupy 2.93mm 2 chip area under 45nm process technology and can run at 800MHz with on average 759.8mW power consumption. ARENA also supports the concurrent execution of multi-applications, offering ideal architectural support for future high-performance parallel computing and data analytics systems.

97 MATHEMATICS AND COMPUTING↗

Width-Based Discharge Partitioning in Distributary Networks: How Right We Are

River deltas are home to large populations and can be composed of complex channel networks which convey flows of matter to the shoreline. Knowledge of flow within individual channels is needed to quantify the distribution of discharge across the delta, and thus its sustainability over time. Due to a lack of field measurements at the local channel scale, researchers leverage remote sensing data to estimate the partitioning of flow. We compare data from 15 river deltas to discharge partitioning estimates based on channel network graphs derived from remote sensing imagery. We quantify errors in the common width-based method and test alternative partitioning techniques to find that width-based discharge partitioning is universally applicable, suggesting that absent any site-specific information, discharge partitioning by average channel width is an appropriate approach. We also provide networks, streamflow measurements, and flux partitioning estimates for 28 delta networks as the Discharge In Distributary NeTworks (DIDNT) dataset.

58 GEOSCIENCES↗

A Semi-Automated Approach for Curating a Glossary of Key Terms for Open-Source Data Queries

In FY20, the Savannah River National Laboratory (SRNL) was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development (NA-22) to build a machine learning based modeling pipeline that could extract proliferation events of interest from open text-based data sources. As a test case, the research team targeted the identification/fusion of events and indicators that fissile core fabrication would be executed at the Savannah River Site prior to its official announcement in May of 2018. The demonstration prototype proved successful by applying natural language processing and graph theoretical techniques to identify contextual shifts in key words and phrases that acted as indicators that pit production would be carried out at the Savannah River Site up to two years prior to the official announcement.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The OEOP Duties of Reasonable Accommodation

I was fortunate enough to be assigned two assignments during my ten weeks here at NASA's Langley Research Center, in the Office of Equal Opportunity Programs (OEOP). One of my projects gave me the chance to gain experience in developing calculation formulas for the EXCEL computer system, while my second project gave me the chance to put my research skills and legal knowledge to use. The function of the OEOP is to ensure the adherence to personnel policy and practices in the employment, development, advancement and treatment of Federal employees and applicants for employment. This includes veterans and disabled as well. My initial project involved the research of hiring and promotion among the different minorities and females employed here at Langley. The objective of my first project was to develop graphs that showed the number of promotions during the past five years for each minority group here on the Center. I also had to show the average number of years it took for each promotion. The objective of my second and main research project was to find and research cases regarding the reasonable accommodation of disabled workers. The research of these cases is to ensure that individuals with disabilities are provided the necessary accommodations that are essential to the function of their job.

Coppedge, Angela↗

Reanalysis of Rodent Data from Spacelab Life Sciences-1

The space bioscience field has long been plagued by the challenge of spaceflight with effects of radiation and microgravity. Having multiple and repeated spaceflight experiments for model organisms to solve these space stressors is costly and time consuming. Therefore, reusing and reanalyzing legacy experiments is one way that scientists can draw new conclusions in a timely manner and without using too many resources. Moreover, advances in general biological knowledge allows legacy experiments to be placed into more complete context.Here we aim to analyze all data and metadata taken from rats flown on the SLS-1 mission to create a comprehensive biological model that can be supplemented with current data to allow new discoveries in how space flown organisms adapt to the space environment. Our approach begins with the identification of all the data and metadata, including graphs and tables, for SLS-1 in NASA archives and other sources. Then, each piece of data and metadata will be digitized, reformatted and analyzed. Lastly, a previously developed astronaut model will be used to create the data framework and a comprehensive biological rodent model. The datasets we are using is from the 1991 SpaceLab Life Science 1 (SLS-1) NASA Mission. This was the first designated spacelab mission flown. All 29 rodents were tested for nine days in two different habitats: Research Animal Holding Facility (RAHF) and Animal Enclosure Module (AEM). The rodents were prepared for a live return and compared to a ground control. A total of 30 rodent experiments were accepted as flight studies on the mission. By digitization and reorganizing SLS-1 rat data we will both directly generate new insights and indirectly enable other scientists to by providing the data and metadata in a digitized form.

Space Biology↗

Reanalysis of Rodent Data from Spacelab Life Science-1

The space bioscience field has long been plagued by the challenge of spaceflight with effects of radiation and microgravity. Having multiple and repeated spaceflight experiments for model organisms to solve these space stressors is costly and time consuming. Therefore, reusing and reanalyzing legacy experiments is one way that scientists can draw new conclusions in a timely manner and without using too many resources. Moreover, advances in general biological knowledge allows legacy experiments to be placed into more complete context.Here we aim to analyze all data and metadata taken from rats flown on the SLS-1 mission to create a comprehensive biological model that can be supplemented with current data to allow new discoveries in how space flown organisms adapt to the space environment. Our approach begins with the identification of all the data and metadata, including graphs and tables, for SLS-1 in NASA archives and other sources. Then, each piece of data and metadata will be digitized, reformatted and analyzed. Lastly, a previously developed astronaut model will be used to create the data framework and a comprehensive biological rodent model. The datasets we are using is from the 1991 SpaceLab Life Science 1 (SLS-1) NASA Mission. This was the first designated spacelab mission flown. All 29 rodents were tested for nine days in two different habitats: Research Animal Holding Facility (RAHF) and Animal Enclosure Module (AEM). The rodents were prepared for a live return and compared to a ground control. A total of 30 rodent experiments were accepted as flight studies on the mission. By digitization and reorganizing SLS-1 rat data we will both directly generate new insights and indirectly enable other scientists to by providing the data and metadata in a digitized form.

Space Biology↗

Figure Descriptive Text Extraction Using Ontological Representation

Abstract Experimental research publications provide figure form resources including graphs, charts, and any type of images to effectively support and convey methods and results. To describe figures, authors add captions, which are often incomplete, and more descriptions reside in body text. This work presents a method to extract figure descriptive text from the body of scientific articles. We adopted ontological semantics to aid concept recognition of figure-related information, which generates human- and machine-readable knowledge representations from sentences. Our results show that conceptual models bring an improvement in figure descriptive sentence classification over word-based approaches.

97 MATHEMATICS AND COMPUTING↗

Automated descriptor selection, volcano curve generation, and active site determination using the DescMAP software

The material space for catalyst discovery is expansive. Volcano curves are traditionally employed to provide physical insights into optimal catalyst characteristics for new material selection. Their generation lies on a single descriptor picked using expert knowledge. Here we present DescMAP, a Python-based software, to automate the selection of descriptors, the generation of volcano maps, and the identification of active sites for structure-sensitive reactions. Here, we consider traditional energy-based and geometric descriptors for structure-sensitive reactions. DescMAP is integrated with the Virtual Kinetic Laboratory (VLab) to provide multiple functionalities. It inputs spreadsheets or template files for flexibility and outputs interactive graphs for post-processing. We demonstrate its features using the non-oxidative dehydrogenation of ethane to ethylene over (111) closed-packed surfaces and the methane total oxidation over various Pt facets. It can be easily applied to other complex chemistries and achieves quick screening of potential catalysts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GRAPH — an readout ASIC for large MCP based detectors

We present a programmable 16 channel, mixed signal, low power readout ASIC, having the project historically named Gigasample Recorder of Analog waveforms from a PHotodetector (GRAPH). It is designed to read large aperture single photon imaging detectors using micro channel plates for charge multiplication, and measuring the detector's response on crossed strips anodes to extrapolate the incoming photon position. Each channel consists of a fast, low power and low noise charge sensitive amplifier, which provides a myriad of coarse and fine programmable options for gain and shaping settings. Further, the amplified signal is recorded using, to our knowledge novel, the Hybrid Universal sampLing Architecture (HULA) ADC. A kind of mixed signal double buffer memory, that enables concurrent waveform recording, and selected event digitized data extraction. The sampling frequency is freely adjustable between few kHz up to 125 MHz, while the chip's internal digital memory holds a history 2048 samples for each channel, with a digital headroom of 12 bits. An optimized region of interest sample-read algorithm allows to extract the information just around the event pulse peak, while selecting the next event, thus substantially reducing the operational dead time. The chip is designed in 130 nm TSMC CMOS technology, and its power consumption is around 47 mW per channel.

47 OTHER INSTRUMENTATION↗

AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis

Checkpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 13 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint.

HPC↗

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Document Retrieval and Ranking using Similarity Graph Mean Hitting Times

We present a novel approach to information retrieval and document analysis based on graph analytic methods. Traditional information retrieval methods use a set of terms to define a query that is applied against a document corpus to identify the documents most related to those terms. In contrast, we define a query as a set of documents of interest and apply the query by computing mean hitting times between this set and all other documents on a document similarity graph abstraction of the semantic relationships between all pairs of documents. We present the steps of our approach along with a simple example application illustrating how this approach can be used to find documents related to two or more documents or topics of interest.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots

Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.

Science & Technology - Other Topics↗

Planetary science: A lunar perspective

An interpretative synthesis of current knowledge on the moon and the terrestrial planets is presented, emphasizing the impact of recent lunar research (using Apollo data and samples) on theories of planetary morphology and evolution. Chapters are included on the exploration of the solar system; geology and stratigraphy; meteorite impacts, craters, and multiring basins; planetary surfaces; planetary crusts; basaltic volcanism; planetary interiors; the chemical composition of the planets; the origin and evolution of the moon and planets; and the significance of lunar and planetary exploration. Photographs, drawings, graphs, tables of quantitative data, and a glossary are provided.

Taylor, S. R.↗

Kinetic parameters from thermogravimetric analysis

High performance polymeric materials are finding increased use in aerospace applications. Proposed high speed aircraft will require materials to withstand high temperatures in an oxidative atmosphere for long periods of time. It is essential that accurate estimates be made of the performance of these materials at the given conditions of temperature and time. Temperatures of 350 F (177 C) and times of 60,000 to 100,000 hours are anticipated. In order to survey a large number of high performance polymeric materials on a reasonable time scale, some form of accelerated testing must be performed. A knowledge of the rate of a process can be used to predict the lifetime of that process. Thermogravimetric analysis (TGA) has frequently been used to determine kinetic information for degradation reactions in polymeric materials. Flynn and Wall studied a number of methods for using TGA experiments to determine kinetic information in polymer reactions. Kinetic parameters, such as the apparent activation energy and the frequency factor, can be determined in such experiments. Recently, researchers at the McDonnell Douglas Research Laboratory suggested that a graph of the logarithm of the frequency factor against the apparent activation energy can be used to predict long-term thermo-oxidative stability for polymeric materials. Such a graph has been called a kinetic map. In this study, thermogravimetric analyses were performed in air to study the thermo-oxidative degradation of several high performance polymers and to plot their kinetic parameters on a kinetic map.

Kiefer, Richard L.↗