Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

143 records · Page 8

cuTS: Scaling Subgraph Isomorphism on Distributed Multi-GPUSystems Using Trie Based Data Structure

Subgraph isomorphism is a pattern-matching algorithm widely used in many domains such as chem-informatics, bioinformatics, databases, and social network analysis. It is computationally expensive and is a proven NP-hard problem. The massive parallelism offered by the GPU hardware is well suited for solving the subgraph isomorphism. However, current GPU implementations are far from the achievable performance. Moreover, the enormous memory requirement of current approaches limits the problem size that can be handled. This work analyzes the fundamental challenges associated with processing the subgraph isomorphism on GPUs and develops an efficient GPU hardware-aware implementation. We also develop a new GPU-friendly trie-based data structure to drastically reduce the intermediate storage space requirement. Hence, our approach runs larger benchmarks than the competitors. We also develop the first distributed sub-graph isomorphism algorithm for GPUs. Our experimental evaluation section demonstrates the efficacy of our approach by comparing the execution time and number of cases that we can handle against the state-of-the-art GPU implementations.

Xiang, Lizhi↗

ChemoGraph: Interactive Visual Exploration of the Chemical Space

Exploratory analysis of the chemical space is an important task in the field of cheminformatics. For example, in drug discovery research, chemists investigate sets of thousands of chemical compounds in order to identify novel yet structurally similar synthetic compounds to replace natural products. Manually exploring the chemical space inhabited by all possible molecules and chemical compounds is impractical, and therefore presents a challenge. To fill this gap, we present ChemoGraph, a novel visual analytics technique for interactively exploring related chemicals. In ChemoGraph, we formalize a chemical space as a hypergraph and apply novel machine learning models to compute related chemical compounds. It uses a database to find related compounds from a known space and a machine learning model to generate new ones, which helps enlarge the known space. Moreover, ChemoGraph highlights interactive features that support users in viewing, comparing, and organizing computationally identified related chemicals. With a drug discovery usage scenario and initial expert feedback from a case study, we demonstrate the usefulness of ChemoGraph.

chemical space exploration↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Can a deep-learning model make fast predictions of vacancy formation in diverse materials?

The presence of point defects, such as vacancies, plays an important role in materials design. Here, we explore the extrapolative power of a graph neural network (GNN) to predict vacancy formation energies. We show that a model trained only on perfect materials can also be used to predict vacancy formation energies (E vac ) of defect structures without the need for additional training data. Such GNN-based predictions are considerably faster than density functional theory (DFT) calculations and show potential as a quick pre-screening tool for defect systems. To test this strategy, we developed a DFT dataset of 530 E vac consisting of 3D elemental solids, alloys, oxides, semiconductors, and 2D monolayer materials. We analyzed and discussed the applicability of such direct and fast predictions. We applied the model to predict 192 494 E vac for 55 723 materials in the JARVIS-DFT database. Our work demonstrates how a GNN-model performs on unseen data.

2D materials↗

Optimizing FPGA-based Accelerator Design for Large-Scale Molecular Similarity Search (Special Session Paper)

Molecular similarity search has been widely used in drug discovery to rapidly identify structurally similar compounds from large molecular databases. With the increasing size of chemical libraries, there is growing interest in the efficient ac- celeration of large-scale similarity search. Existing works mainly focus on CPU and GPU to accelerate the computation of Tatimoto coefficient in measuring the pairwise similarity between different molecular fingerprints. In this paper, we propose and optimize an FPGA-based accelerator design on exhaustive and approximate search algorithms. On exhaustive search using BitBound & fold- ing, we analyze the similarity cutoff and folding level relationship with search speedup and accuracy, and propose a scalable on- the-fly query engine on FPGAs to reduce the resource utilization and pipeline interval. We achieve a 450 million compounds-per- second processing throughput for a single query engine. On approximate search using hierarchical navigable small world (HNSW), a popular algorithm with high recall and query speed, we propose an FPGA-based graph traversal engine to utilize high throughput register array based priority queue and fine- grained distance calculation engine to increase the processing capability. Experimental results show that the proposed FPGA- based HNSW implementation achieves a 35× speedup than existing works on CPU. To the best of our knowledge, our FPGA- based implementation is the first attempt to accelerate molecular similarity search on FPGA and has the highest performance among existing approaches.

Peng, Hongwu↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry↗

Influence of Thermophysical Property Variability on Thermal Predictions of Ti-6Al-4V

The additive manufacturing (AM) industry has experienced rapid growth in recent decades as industrial interest in the process has grown. Because of this, there are many groups interested in simulations of the AM process. However, the influence of thermophysical property variability on the melt pool geometry during processing of various AM alloys is unclear. The goal of my work at NASA Langley Research Center (LaRC) was to characterize this influence on Ti-6Al-4V during AM processing alongside developing a tool to enable equivalent studies on other AM materials. To facilitate this process, a database tool was developed to contain and organize thermophysical data previously reported by primary sources. The specific thermophysical properties of interest for this study were density, specific heat capacity, and conductivity. Thermal diffusivity was calculated using the other three properties. This data was then fit with a Gaussian distribution and then randomly sampled from using Monte Carlo random value sampling. Using the Rosenthal equation, this sample data was used to simulate the temperature field during additive manufacturing and extract the melt pool geometry. Specifically, the melt pool’s length, width, and depth. This process was repeated an arbitrary number of times, with the default being 1000. Once this information was obtained, histograms were made showing the distributions of the sizes of the melt pool’s length, width, and depth. Representative statistical metrics of the distributions were calculated (mean, standard deviation, and coefficient of variance). This work found that at the simulated processing values, the melt pool’s width and depth had a 5.7% variation while the length had a 3.2% variation. Additionally, a tool was constructed to graph a thermal color map representation of the melt pool for specific, arbitrary values of density, specific heat, and conductivity. This tool allowed for more efficient plotting of individual queries of the thermophysical properties. The results that were observed in the research were that values reported in literature vary and this variance can have a significant impact on the simulation. Predictions of the melt pool’s dimensions show that length has a standard deviation of approximately 6 μm, width has a standard deviation of approximately 5 μm, and depth can vary by approximately 3 μm. Considering the mean sizes of length, width, and depth are 185 μm, 88 μm, and 44 μm respectively, such a deviation is significant. This shows that depth, for example, could be more than 10% larger or smaller than expected. This demonstrates that variability in the reported thermophysical properties are not negligible and should be expected to have an influence on the results of the laser powder bed fusion additive manufacturing process. When simulating this process in the future, measures should be taken to account for this discrepancy and the uncertainty involved

Justin Martin↗

Advanced Technology Lifecycle Analysis System (ATLAS)

Developing credible mass and cost estimates for space exploration and development architectures require multidisciplinary analysis based on physics calculations, and parametric estimates derived from historical systems. Within the National Aeronautics and Space Administration (NASA), concurrent engineering environment (CEE) activities integrate discipline oriented analysis tools through a computer network and accumulate the results of a multidisciplinary analysis team via a centralized database or spreadsheet Each minute of a design and analysis study within a concurrent engineering environment is expensive due the size of the team and supporting equipment The Advanced Technology Lifecycle Analysis System (ATLAS) reduces the cost of architecture analysis by capturing the knowledge of discipline experts into system oriented spreadsheet models. A framework with a user interface presents a library of system models to an architecture analyst. The analyst selects models of launchers, in-space transportation systems, and excursion vehicles, as well as space and surface infrastructure such as propellant depots, habitats, and solar power satellites. After assembling the architecture from the selected models, the analyst can create a campaign comprised of missions spanning several years. The ATLAS controller passes analyst specified parameters to the models and data among the models. An integrator workbook calls a history based parametric analysis cost model to determine the costs. Also, the integrator estimates the flight rates, launched masses, and architecture benefits over the years of the campaign. An accumulator workbook presents the analytical results in a series of bar graphs. In no way does ATLAS compete with a CEE; instead, ATLAS complements a CEE by ensuring that the time of the experts is well spent Using ATLAS, an architecture analyst can perform technology sensitivity analysis, study many scenarios, and see the impact of design decisions. When the analyst is satisfied with the system configurations, technology portfolios, and deployment strategies, he or she can present the concepts to a team, which will conduct a detailed, discipline-oriented analysis within a CEE. An analog to this approach is the music industry where a songwriter creates the lyrics and music before entering a recording studio.

O'Neil, Daniel A.↗

Automated Processing of ISIS Topside Ionograms into Electron Density Profiles

Modeling of the topside ionosphere has for the most part relied on just a few years of data from topside sounder satellites. The widely used Bent et al. (1972) model, for example, is based on only 50,000 Alouette 1 profiles. The International Reference Ionosphere (IRI) (Bilitza, 1990, 2001) uses an analytical description of the graphs and tables provided by Bent et al. (1972). The Alouette 1, 2 and ISIS 1, 2 topside sounder satellites of the sixties and seventies were ahead of their times in terms of the sheer volume of data obtained and in terms of the computer and software requirements for data analysis. As a result, only a small percentage of the collected topside ionograms was converted into electron density profiles. Recently, a NASA-funded data restoration project has undertaken and is continuing the process of digitizing the Alouette/ISIS ionograms from the analog 7-track tapes. Our project involves the automated processing of these digital ionograms into electron density profiles. The project accomplished a set of important goals that will have a major impact on understanding and modeling of the topside ionosphere: (1) The TOPside Ionogram Scaling and True height inversion (TOPIST) software was developed for the automated scaling and inversion of topside ionograms. (2) The TOPIST software was applied to the over 300,000 ISIS-2 topside ionograms that had been digitized in the fkamework of a separate AISRP project (PI: R.F. Benson). (3) The new TOPIST-produced database of global electron density profiles for the topside ionosphere were made publicly available through NASA s National Space Science Data Center (NSSDC) ftp archive at . (4) Earlier Alouette 1,2 and ISIS 1, 2 data sets of electron density profiles from manual scaling of selected sets of ionograms were converted fiom a highly-compressed binary format into a user-friendly ASCII format and made publicly available through nssdcftp.gsfc.nasa.gov. The new database for the topside ionosphere established as a result of this project, has stimulated a multitude of new studies directed towards a better description and prediction of the topside ionosphere. Marinov et al. (2004) developed a new model for the upper ion transition height (Oxygen to Hydrogen and Helium) and Bilitza (2004) deduced a correction term for the I N topside electron density model. Kutiev et al. (2005) used this data to develop a new model for the topside ionosphere scale height (TISH) as a function of month, local time, latitude, longitude and solar flux F10.7. Comparisons by Belehaki et al. (2005) show that TISH is in general agreement with scale heights deduced from ground ionosondes but the model predicts post-midnight and afternoon maxima whereas the ionosonde data show a noon maximum. Webb and Benson (2005) reported on their effort to deduce changes in the plasma temperature and ion composition from changes in the topside electron density profile as recorded by topside sounders. Limitations and possible improvements of the IRI topside model were discussed by Coisson et al. (2005) including also the possible use of the NeQuick model, Our project progressed in close collaboration and coordination with the GSFC team involved in the ISIS digitization effort. The digitization project was highly successful producing a large amount of digital topside ionograms. Several no-cost extensions of the TOPIST project were necessary to keep up with the pace and volume of the digitization effort.

Reinisch, bodo W.↗

GROWdb US River Systems - Samples

GROW Overview We developed the Genome Resolved Open Watersheds database (GROWdb), which aims to increase genomic sampling and understanding of global river microbiomes. An emphasis of GROWdb is to create a publicly available and ever-expanding microbial genome database that is focused on rivers while being interoperable with databases from other ecosystems. GROWdb is based on a network-of-networks approach to move beyond a small collection of well-studied rivers, towards a spatially distributed, global network of systematic observations. GROWdb represents the first microbial, river-focused resource parsed at various scales from genes to MAGs to community level including expression and potential based measurements that will be of interest to microbiologists, ecologists, geochemists, hydrologists, and modelers. Dataset Acknowledgement GROWdb contains data from various research campaigns, please acknowledge the following data generators, as appropriate: WHONDRS derived genomes or samples - include this statement in your acknowledgements: “This study used data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) under the River Corridor Science Focus Area (SFA) at the Pacific Northwest National Laboratory (PNNL) that was generated at the U.S. Department of Energy (DOE) Joint Genome Institute User Facility. PNNL is operated by Battelle Memorial Institute for the U.S. DOE under Contract No. DE-AC05-76RL01830. The SFA is supported by the U.S. DOE, Office of Biological and Environmental Research (BER), Environmental System Science (ESS) Program.” Total Samples loaded onto this Narrative: 178 Note: Not all GROW samples may be loaded into KBase Data Availability The data underlying GROWdb are accessible across various platforms to ensure all levels of data structure are widely available. First, all reads and MAGs are publicly hosted on National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. Second, all data related data presented here including MAG annotations, extended data tables, phylogenetic tree files, antibiotic resistance gene database files, and MAG abundance tables are available in Zenodo (link). Beyond the flat database files listed above, our aim for GROWdb was to maximize data use by making the data available in searchable and interactive platforms including the National Microbiome Data Collaborative (NMDC) data portal, the Department of Energy’s Systems Biology Knowledgebase (KBase), and a GROW specific user interface released here, GROWdb Explorer. Each platform provides different ways to interact with GROWdb: NMDC GROWdb formed a pilot project for the NMDC. Specifically, individual GROWdb datasets (metagenomes, metatranscriptomes, etc) are easily accessible and searchable through the NMDC data portal, where they are systematically connected to each other and to a rich suite of sample information and standard analysis results, following Findable, Accessible, Interoperable, and Reusable (FAIR) data practices. KBase GROWdb is publicly available within KBase, including samples (this Narrative), MAGs, and corresponding genome scale metabolic models. Access within KBase allows for immediate access and reuse of data, including comparison to private data using KBase’s 500+ analysis tools. Other linked narratives in KBase: GROW Metagenome Assembled Genomes (MAGs) GROW Metabolic Models GROWdb Explorer GROWdb data is also explorable through a graphical user interface built through the Colorado State University Geospatial Centroid (https://geocentroid.shinyapps.io/GROWdatabase/), allowing users to search and graph microbial and spatial data simultaneously. In summary, this microbial genome resource represents the first publicly available genome collection from rivers and offers data that can be leveraged across microbiome studies. GROWdb is an expanding repository to incorporate and unify global river multi-omic data for the future.

59 BASIC BIOLOGICAL SCIENCES↗

Automated ISS Flight Utilities

During my internship at NASA Johnson Space Center, I worked in the Space Radiation Analysis Group (SRAG), where I was tasked with a number of projects focused on the automation of tasks and activities related to the operation of the International Space Station (ISS). As I worked on a number of projects, I have written short sections below to give a description for each, followed by more general remarks on the internship experience. My first project is titled "General Exposure Representation EVADOSE", also known as "GEnEVADOSE". This project involved the design and development of a C++/ ROOT framework focused on radiation exposure for extravehicular activity (EVA) planning for the ISS. The utility helps mission managers plan EVAs by displaying information on the cumulative radiation doses that crew will receive during an EVA as a function of the egress time and duration of the activity. SRAG uses a utility called EVADOSE, employing a model of the space radiation environment in low Earth orbit to predict these doses, as while outside the ISS the astronauts will have less shielding from charged particles such as electrons and protons. However, EVADOSE output is cumbersome to work with, and prior to GEnEVADOSE, querying data and producing graphs of ISS trajectories and cumulative doses versus egress time required manual work in Microsoft Excel. GEnEVADOSE automates all this work, reading in EVADOSE output file(s) along with a plaintext file input by the user providing input parameters. GEnEVADOSE will output a text file containing all the necessary dosimetry for each proposed EVA egress time, for each specified EVADOSE file. It also plots cumulative dose versus egress time and the ISS trajectory, and displays all of this information in an auto-generated presentation made in LaTeX. New features have also been added, such as best-case scenarios (egress times corresponding to the least dose), interpolated curves for trajectories, and the ability to query any time in the EVADES output. As mentioned above, GEnEVADOSE makes extensive use of ROOT version 6, the data analysis framework developed at the European Organization for Nuclear Research (CERN), and the code is written to the C++11 standard (as are the other projects). My second project is the Automated Mission Reference Exposure Utility (AMREU).Unlike GEnEVADOSE, AMREU is a combination of three frameworks written in both Python and C++, also making use of ROOT (and PyROOT). Run as a combination of daily and weekly cron jobs, these macros query the SRAG database system to determine the active ISS missions, and query minute-by-minute radiation dose information from ISS-TEPC (Tissue Equivalent Proportional Counter), one of the radiation detectors onboard the ISS. Using this information, AMREU creates a corrected data set of daily radiation doses, addressing situations where TEPC may be offline or locked up by correcting doses for days with less than 95% live time (the total amount time the instrument acquires data) by averaging the past 7 days. As not all errors may be automatically detectable, AMREU also allows for manual corrections, checking an updated plaintext file each time it runs. With the corrected data, AMREU generates cumulative dose plots for each mission, and uses a Python script to generate a flight note file (.docx format) containing these plots, as well as information sections to be filled in and modified by the space weather environment officers with information specific to the week. AMREU is set up to run without requiring any user input, and it automatically archives old flight notes and information files for missions that are no longer active. My other projects involve cleaning up a large data set from the Charged Particle Directional Spectrometer (CPDS), joining together many different data sets in order to clean up information in SRAG SQL databases, and developing other automated utilities for displaying information on active solar regions, that may be used by the space weather environment officers to monitor solar activity. I consulted my mentor Dr. Ryan Rios and Dr. Kerry Lee for project requirements and added features, and ROOT developer Edmond Offermann for advice on using the ROOT library. I also received advice and feedback from Dr. Janet Barzilla of SRAG, who tested my code. Besides these inputs, I worked independently, writing all of the code by myself. The code for all these projects is documented throughout, and I have attempted to write it in a modular format. Assuming that ROOT is updated accordingly, these codes are also Y2038-compliant (and Y10K-compliant). This allows the code to be easily referenced, modified and possibly repurposed for non-ISS missions in the future, should the necessary inputs exist. These projects have taught me a lot about coding and software design - I have become a much more skilled C++ programmer and ROOT user, and I also learned to code in Python and PyROOT (and its advantages and disadvantages compared to C++/ ROOT). Furthermore, I have learned about space radiation and radiation modeling, topics that greatly interest me as I pursue a degree in physics. Working alongside experimental physicists like Dr. Rios, I have developed a greater understanding and appreciation for experimental science, something I have always leaned towards but to which I lacked significant exposure. My work in SRAG has also given me the invaluable opportunity to witness the work environment for physicists at NASA, and what a career in academia may look like at a government laboratory such as NASA Johnson Space Center. As I continue my studies and look forward to graduate school and a future career, this experience at NASA has given me a meaningful and enjoyable opportunity to put my skills to use and see what my future career path might hold.

Offermann, Jan Tuzlic↗

Long Duration Exposure Facility (LDEF) experiment M0003 meteoroid and debris survey

A survey of the meteoroid and space debris impacts on LDEF experiment M0003 was performed. The purpose of this survey was to document significant impact phenomenology and to obtain impact crater data for comparison to current space debris and micrometeoroid models. The survey consists of the following: photomicrographs of significant impacts in a variety of material types; accurate measurements of impact crater coordinates and dimensions for selected experiment surfaces; and databasing of the crater data for reduction, manipulation, and comparison to models. Large area surfaces that were studied include the experiment power and data system (EPDS) sunshields, environment exposure control canister (EECC) sunshields, and the M0003 signal conditioning unit (SCU) covers. Crater diameters down to 25 microns were measured and cataloged. Both leading (D8) and trailing (D4) edge surfaces were studied and compared. The EPDS sunshields are aluminum panels painted with Chemglaze A-276 white thermal control paint, the EECC sunshields are chromic acid-anodized aluminum, and the SCU covers are aluminum painted with S13GLO white thermal control paint. Typical materials that have documented impacts are metals, glasses and ceramics, composites, polymers, electronic materials, and paints. The results of this survey demonstrate the different response of materials to hypervelocity impacts. Comparison of the survey data to curves derived from the Kessler debris model and the Cour-Palais micrometeoroid model indicates that these models overpredict small impacts (less than 100 micron) and may underpredict large impacts (greater than 1000 micron) while having fair to good agreement for the intermediate impacts. Comparison of the impact distributions among the various surfaces indicates significant variations, which may be a function of material response effects, or in some cases surface roughness. Representative photographs and summary graphs of the impact data are presented.

Meshishnek, M. J.↗

An experimental study of the sources of fluctuating pressure loads beneath swept shock/boundary-layer interactions

An experimental research program providing basic knowledge and establishing a database on the fluctuating pressure loads produced on aerodynamic surfaces beneath three dimensional shock wave/boundary layer interactions is described. Such loads constitute a fundamental problem of critical concern to future supersonic and hypersonic flight vehicles. A turbulent boundary layer on a flat plate is subjected to interactions with swept planar shock waves generated by sharp fins at angle of attack. Fin angles from 10 to 20 deg at freestream Mach numbers of 3 and 4 produce a variety of interaction strengths from weak to very strong. Miniature Kulite pressure transducers flush-mounted in the flat plate are used to measure interaction-induced wall pressure fluctuations. The distributions of properties of the pressure fluctuations, such as their ring levels, amplitude distributions, and power spectra, are also determined. Measurements were made for the first time in the aft regions of these interactions, revealing fluctuating pressure levels as high as 160 dB. These fluctuations are dominated by low frequency (0-5 kHz) signals. The maximum ring levels in the interactions show an increasing trend with increasing interaction strength. On the other hand, the maximum ring levels in the forward portion of the interactions decrease linearly with increasing interaction sweep back. These ring pressure distributions and spectra are correlated with the features of the interaction flowfield. The unsteadiness of the off-surface flowfield is studied using a new, non-intrusive technique based on the shadow graph method. The results indicate that the entire lambda-shock structure generated by the interaction undergoes relatively low-frequency oscillations. Some regions where particularly strong fluctuations are generated were identified. Fluctuating pressure measurements are also made along the line of symmetry of an axisymmetric jet impinging upon a flat plate at an angle. This flow was chosen as a simple analog to the impinging jet region found in the rear portion of the shock wave/boundary layer interactions under study. It is found that a sharp peak in ring pressure level exists at or near the mean stagnation point. It is suggested that the phenomena responsible for this peak may be active in the swept interactions as well, and may cause the extremely high fluctuating pressures observed in the impinging jet region in the present experimental program.

Settles, G. S.↗

Verification Testing of OLI Systems Mixed Solvent Electrolyte Model for the Na-K-Mg-Ca-H-Cl-SO 4 -OH-HCO 3 -CO 3 -CO 2 -H 2 ) System to High Ionic Strength at 25°C

This technical report summarizes model verification results and summary statistics for 41 evaporite mineral solubility cases evaluated by Savannah River National Laboratory using OLI Systems’ aqueous electrolyte thermodynamic modeling software. The 41 verification cases containing a total of 60 solubility curves comprise mineral solubility data from low to high ionic strength at 25°C for the eight-component system Na-K-Mg-Ca-H-Cl-SO 4 -OH-HCO 3 -CO 3 -CO 2 -H 2 O as reported by Harvie et al. (1984). Thermodynamic calculations were executed using OLI Systems’ Stream Analyzer computation module within the OLI Studio software platform (Ver. 11.0, Rev. 11.0.1.9). The Mixed Solvent Electrolyte (MSE) thermodynamic framework was chosen for this investigation because of its superiority in modeling high ionic-strength inorganic salt solutions and actinide redox chemistry and solubility, both of which are relevant to the geological repository conditions at the Waste Isolation Pilot Plant in Carlsbad, New Mexico. Mineral solubility data in various inorganic salt solutions were digitized and extracted from figures generated by Harvie et al. (1984). For each of the 60 solubility curves, a case-specific chemistry model and input file were generated in OLI Studio using OLI Stream Analyzer and the MSE (H 3 O + ion) public databank provided by OLI Systems. Model simulation results were exported to Microsoft Excel to calculate summary statistics and to generate graphs comparing the OLI model predictions to the solubility data. Summary statistics include residuals (model – data) and concordance (accuracy × precision, where precision is indicated by the Pearson correlation coefficient and accuracy accounts for bias and scale differential). Private databanks were not developed, and activity coefficient model regressions were not performed to improve OLI model fits to the data. Of the 41 model verification plots, 83% have a mean of the percent residuals less than or equal to 25%. Similarly, 75% display a concordance greater than or equal to 0.75. Only seven of the 41 verification plots fail to show good agreement between the model and data. Of these seven, three are relevant to the WIPP repository because they involve the Mg-OH-Cl-SO 4 -CO 3 aqueous system. The remaining four address salt solubilities at the pH extremes (strong acid and strong base). It should be noted that in two of the three Mg-OH-Cl-SO 4 -CO 3 system cases, the regressed Harvie et al. (1984) solubility curve also deviated from the data. Lack of agreement between the OLI model-predicted solubility curves and the data is attributable to one or more of the following: specific solid species are not included in the OLI MSE databank; there is significant variation among the different solubility datasets chosen by Harvie et al. (1984); the OLI MSE model’s thermodynamic parameters were determined using different solubility datasets; and the activity coefficient parameters for certain relevant ion-ion and ion-molecule pairs have not been optimized via data regression. Two recommendations for future work are to (1) evaluate solubility data for the Mg-OH-Cl-SO 4 -CO 3 system at high ionic strength and, if necessary, develop a private OLI MSE database that includes missing species and, where necessary, regressed standard state properties and interaction parameters; (2) perform similar verification testing of the OLI model for actinide solubility data.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

LSKnowledge: Nexus for Transformative Scientific Discoveries and Enhanced Information Retrieval in NASA Life Sciences Portal

We stand at the brink of an extraordinary transformation in the field of AI, driven by the convergence of generative AI and semantic technologies (e.g., knowledge graphs). This fusion holds immense potential and could redefine the future of scientific exploration, particularly in the realm of life sciences research. In this context, we shed light on the pivotal roles that Large Language Models (LLMs) and semantic technologies will play in advancing research, unearthing and comprehending life sciences information through innovative approaches, and empowering researchers to extract insights from NASA's extensive Life Sciences Data Archive. Within the NASA Life Sciences Portal (NLSP), the integration of LLMs and semantic technologies unlocks several advanced capabilities. First and foremost, it equips scientists with sophisticated tools to manage the ever-expanding wealth of scientific literature and data. Furthermore, it facilitates the creation of knowledge graphs that visually represent intricate relationships among biological entities, enabling comprehensive systems-level analysis. Additionally, the fusion of generative AI (including LLMs) and semantic technology can significantly benefit NASA's life sciences research by enhancing information retrieval and hypothesis generation. These tools enhance natural language understanding, facilitating knowledge discovery within NLSP. The overarching vision is to establish a cohesive knowledge ecosystem within NLSP, harnessing the power of LLMs and semantic technologies to synthesize and cross-reference data from diverse missions, disciplines, and research domains. This holistic approach ultimately deepens our understanding of how space environments impact life sciences data. To advance this initiative, we have launched LSKnowledge, aimed at enhancing the information retrieval capabilities of NLSP. In the short term, our primary goal is to develop a robust semantic search system. This system will empower HRP (Human Research Program) researchers to navigate NLSP data repositories more efficiently and precisely, catalyzing the process of hypothesis formation and scientific breakthroughs. To achieve this, we have employed pre-trained LLMs as part of a semantic search tool that can rank and highlight the most relevant records for user queries. To assess the tool's performance, we have curated a set of approximately 200 queries from subject matter experts (SMEs) and manually ranked the top records retrieved by both the current search system and the new semantic search, using SME judgments as the gold standard for relevancy. Herein, we present the results of our comparative analysis and illustrate how these findings have informed the fine-tuning of the system for enhanced performance. In the long term, our objectives include 1) retrieving publicly available information and integrating it with NLSP data to provide more precise answers to user queries, and 2) incorporating non-textual information from the NLSP database into our approach. In conclusion, the fusion of LLMs and semantic technologies within NLSP represents a pioneering stride towards reshaping the landscape of scientific discovery. This synergy not only equips researchers with powerful tools to navigate the burgeoning sea of information but also facilitates a deeper understanding of complex biological relationships, all while accelerating hypothesis generation and knowledge discovery. Through our initiative, LSKnowledge, we are committed to continually refining and expanding these capabilities, with the aim of not only enhancing information retrieval but also integrating diverse data sources to provide more precise insights. In the grand vision, NLSP strives to become the cornerstone of a comprehensive knowledge ecosystem, unraveling the enigmatic intricacies of life sciences phenomena in the context of space environments.

Life Sciences↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗