Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Graph processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Hazards, Safety and Design Considerations for Commercial Lithium-ion Cells and Batteries

This viewgraph presentation reviews the features of the Lithium-ion batteries, particularly in reference to the hazards and safety of the battery. Some of the characteristics of the Lithium-ion cell are: Highest Energy Density of Rechargeable Battery Chemistries, No metallic lithium, Leading edge technology, Contains flammable electrolyte, Charge cut-off voltage is critical (overcharge can result in fire), Open circuit voltage higher than metallic lithium anode types with similar organic electrolytes. Intercalation is a process that places small ions in crystal lattice. Small ions (such as lithium, sodium, and the other alkali metals) can fit in the interstitial spaces in a graphite lattice. These metallic ions can go farther and force the graphitic planes apart to fit two, three, or more layers of metallic ions between the carbon sheets. Other features of the battery/cell are: The graphite is conductive, Very high energy density compared to NiMH or NiCd, Corrosion of aluminum occurs very quickly in the presence of air and electrolyte due to the formation of HF from LiPF6 and HF is highly corrosive. Slides showing the Intercalation/Deintercalation and the chemical reactions are shown along with the typical charge/discharge for a cylindrical cell. There are several graphs that review the hazards of the cells.

Jeevarajan, Judith↗

A reusable knowledge acquisition shell: KASH

KASH (Knowledge Acquisition SHell) is proposed to assist a knowledge engineer by providing a set of utilities for constructing knowledge acquisition sessions based on interviewing techniques. The information elicited from domain experts during the sessions is guided by a question dependency graph (QDG). The QDG defined by the knowledge engineer, consists of a series of control questions about the domain that are used to organize the knowledge of an expert. The content information supplies by the expert, in response to the questions, is represented in the form of a concept map. These maps can be constructed in a top-down or bottom-up manner by the QDG and used by KASH to generate the rules for a large class of expert system domains. Additionally, the concept maps can support the representation of temporal knowledge. The high degree of reusability encountered in the QDG and concept maps can vastly reduce the development times and costs associated with producing intelligent decision aids, training programs, and process control functions.

Westphal, Christopher↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗

Incorporation of Human Risk Directed Acyclic Graphs (DAG) With Mishap Investigations to Un-Silo Knowledge

NASA’s Human System Risk Board (HSRB) has been a central driver in efforts to understand, mitigate, and communicate the 29 human systems risks monitored by the board. As a result of the collaboration between research, operations, and technical authorities, large bodies of knowledge have been collected and digested to represent the current understanding of the risks. As a part of these bodies of knowledge, directed acyclic graphs (DAGs) have been developed to communicate the current understanding of the causal relationship of the hazards, contributing factors, countermeasures, other risks, and outcomes that contribute to the overall risk. This risk knowledge is applied in a theoretical sense for potential incidents during exploration even while informed by surveillance data. However, there have been mishaps and close calls during past space exploration that intersect with one or more of the Human System Risks DAGs and knowledge bases. The purpose of this exercise was to un-silo this risk knowledge and connect it to the close call of EVA 23 through the development of a DAG representing the intersection of the HSRB Risks and the events of the close call. The development of the DAG occurred through an iterative process, with each iteration expanding and/or refining the nodes and connections described by the source materials. In addition to the risk documentation developed by the HSRB, lessons learned and other mishap investigation documents were utilized to understand the events that led to water entering the helmet of a crewmember on the EVA. New nodes specific to the events of EVA 23 were interconnected with existing HSRB DAG nodes and edges. Nodes within the DAG were defined within a “DAG-tionary” with any updates to a definition that may have previously existed from the HSRB DAGs, and edges were recorded in a matrix. Both the DAG-tionary and matrix describe where nodes and edges are present across the Risk and Mishap DAG. This DAG will then be reviewed by experts outside of HSRB and HRP to confirm that interpretations of the non-health related events (such as the engineering nodes) are represented accurately. DISCUSSION This process highlighted a method by which the knowledge generated among the contributing members of the HSRB Risks can be effectively adapted and utilized through the tools employed by the Risk Custodian teams. By leveraging these tools, new context and insights to the information at hand can be brought forward to address current spaceflight challenges. Moreover, un-siloing this knowledge through future DAGs and other efforts can drive interprofessional collaboration and foster communication. This will enable teams to work together more effectively, leveraging their diverse expertise to tackle the complex challenges of space exploration and human research. Ultimately, this collaboration will bring NASA closer to achieve agency goals and contribute to the overall shared mission and vision.

Samuel Jacobs↗

New High-Performance SiC Fiber Developed for Ceramic Composites

Sylramic-iBN fiber is a new type of small-diameter (10-mm) SiC fiber that was developed at the NASA Glenn Research Center and was recently given an R&D 100 Award for 2001. It is produced by subjecting commercially available Sylramic (Dow Corning, Midland, MI) SiC fibers, fabrics, or preforms to a specially designed high-temperature treatment in a controlled nitrogen environment for a specific time. It can be used in a variety of applications, but it currently has the greatest advantage as a reinforcement for SiC/SiC ceramic composites that are targeted for long-term structural applications at temperatures higher than the capability of metallic superalloys. The commercial Sylramic SiC fiber, which is the precursor for the Sylramic-iBN fiber, is produced by Dow Corning, Midland, Michigan. It is derived from polymers at low temperatures and then pyrolyzed and sintered at high temperatures using boron-containing sintering aids (ref. 1). The sintering process results in very strong fibers (>3 GPa) that are dense, oxygen-free, and nearly stoichiometric. They also display an optimum grain size that is beneficial for high tensile strength, good creep resistance, and good thermal conductivity (ref. 2). The NASA-developed treatment allows the excess boron in the bulk to diffuse to the fiber surface where it reacts with nitrogen to form an in situ boron nitride (BN) coating on the fiber surface (thus the product name of Sylramic-iBN fiber). The removal of boron from the fiber bulk allows the retention of high tensile strength while significantly improving creep resistance and electrical conductivity, and probably thermal conductivity since the grains are slightly larger and the grain boundaries cleaner (ref. 2). Also, as shown in the graph, these improvements allow the fiber to display the best rupture strength at high temperatures in air for any available SiC fiber. In addition, for CMC applications under oxidizing conditions, the formation of an in situ BN surface layer creates a more environmentally durable fiber surface not only because a more oxidation-resistant BN is formed, but also because this layer provides a physical barrier between contacting fibers with oxidation-prone SiC surface layers (refs. 3 and 4). This year, Glenn demonstrated that the in situ BN treatment can be applied simply to Sylramic fibers located within continuous multifiber tows, within woven fabric pieces, or even assembled into complex product shapes (preforms). SiC/SiC ceramic composite panels have been fabricated from Sylramic-iBN fabric and then tested at Glenn within the Ultra-Efficient Engine Technology Program. The test conditions were selected to simulate those experienced by hot-section components in advanced gas turbine engines. The results from testing at Glenn demonstrate all the benefits expected for the Sylramic-iBN fibers. That is, the composites displayed the best thermostructural performance in comparison to composites reinforced by Sylramic fibers and by all other currently available high-performance SiC fiber types (refs. 3 and 5). For these reasons, the Ultra-Efficient Engine Technology Program has selected the Sylramic-iBN fiber for ongoing efforts aimed at SiC/SiC engine component development.

DiCarlo, James A.↗

Representation and display of vector field topology in fluid flow data sets

The visualization of physical processes in general and of vector fields in particular is discussed. An approach to visualizing flow topology that is based on the physics and mathematics underlying the physical phenomenon is presented. It involves determining critical points in the flow where the velocity vector vanishes. The critical points, connected by principal lines or planes, determine the topology of the flow. The complexity of the data is reduced without sacrificing the quantitative nature of the data set. By reducing the original vector field to a set of critical points and their connections, a representation of the topology of a two-dimensional vector field that is much smaller than the original data set but retains with full precision the information pertinent to the flow topology is obtained. This representation can be displayed as a set of points and tangent curves or as a graph. Analysis (including algorithms), display, interaction, and implementation aspects are discussed.

Helman, James↗

Wavelet Methods Developed to Detect and Control Compressor Stall

A "wavelet" is, by definition, an amplitude-varying, short waveform with a finite bandwidth (e.g., that shown in the first two graphs). Naturally, wavelets are more effective than the sinusoids of Fourier analysis for matching and reconstructing signal features. In wavelet transformation and inversion, all transient or periodic data features (as in compressor-inlet pressures) can be detected and reconstructed by stretching or contracting a single wavelet to generate the matching building blocks. Consequently, wavelet analysis provides many flexible and effective ways to reduce noise and extract signals which surpass classical techniques - making it very attractive for data analysis, modeling, and active control of stall and surge in high-speed turbojet compressors. Therefore, fast and practical wavelet methods are being developed in-house at the NASA Lewis Research Center to assist in these tasks. This includes establishing user-friendly links between some fundamental wavelet analysis ideas and the classical theories (or practices) of system identification, data analysis, and processing.

Le, Dzu K.↗

Pulse Detonation Engine Modeled

Pulse Detonation Engine Technology is currently being investigated at Glenn for both airbreathing and rocket propulsion applications. The potential for both mechanical simplicity and high efficiency due to the inherent near-constant-volume combustion process, may make Pulse Detonation Engines (PDE's) well suited for a number of mission profiles. Assessment of PDE cycles requires a simulation capability that is both fast and accurate. It should capture the essential physics of the system, yet run at speeds that allow parametric analysis. A quasi-one-dimensional, computational-fluid-dynamics-based simulation has been developed that may meet these requirements. The Euler equations of mass, momentum, and energy have been used along with a single reactive species transport equation, and submodels to account for dominant loss mechanisms (e.g., viscous losses, heat transfer, and valving) to successfully simulate PDE cycles. A high-resolution numerical integration scheme was chosen to capture the discontinuities associated with detonation, and robust boundary condition procedures were incorporated to accommodate flow reversals that may arise during a given cycle. The accompanying graphs compare experimentally measured and computed performance over a range of operating conditions for a particular PDE. Experimental data were supplied by Fred Schauer and Jeff Stutrud from the Air Force Research Laboratory at Wright-Patterson AFB and by Royce Bradley from Innovative Scientific Solutions, Inc. The left graph shows thrust and specific impulse, Isp, as functions of equivalence ratio for a PDE cycle in which the tube is completely filled with a detonable hydrogen/air mixture. The right graph shows thrust and specific impulse as functions of the fraction of the tube that is filled with a stoichiometric mixture of hydrogen and air. For both figures, the operating frequency was 16 Hz. The agreement between measured and computed values is quite good, both in terms of trend and magnitude. The error is under 10 percent everywhere except for the thrust value at an equivalence ratio of 0.8 in the left figure, where it is 14 percent. The simulation results shown were made using 200 numerical cells. Each cycle of the engine, approximately 0.06 sec, required 2.0 min of CPU time on a Sun Ultra2. The simulation is currently being used to analyze existing experiments, design new experiments, and predict performance in propulsion concepts where the PDE is a component (e.g., hybrid engines and combined cycles).

Paxson, Daniel E.↗

Planar Particle Imaging Doppler Velocimetry Developed

Two current techniques exist for the measurement of planar, three-component velocity fields. Both techniques require multiple views of the illumination plane in order to extract all three velocity components. Particle image velocimetry (PIV) is a high-resolution, high accuracy, planar velocimetry technique that provides valuable instantaneous velocity information in aeropropulsion test facilities. PIV can provide three-component flow-field measurements using a two-camera, stereo viewing configuration. Doppler global velocimetry (DGV) is another planar velocimetry technique that can provide three component flow-field measurements; however, it requires three detector systems that must be located at oblique angles from the measurement plane. The three-dimensional configurations of either technique require multiple (DGV) or at least large (stereo PIV) optical access ports in the facility in which the measurements are being conducted. Optical access is extremely limited in aeropropulsion test facilities. In many cases, only one optical access port is available. A hybrid measurement technique has been developed at the NASA Glenn Research Center, planar particle image and Doppler velocimetry (PPIDV), which combines elements from both the PIV and DGV techniques into a single detection system that can measure all three components of velocity across a planar region of a flow field through a single optical access port. In the standard PIV technique, a pulsed laser is used to illuminate the flow field at two closely spaced instances in time, which are recorded on a "frame-straddling" camera, yielding a pair of single-exposure image frames. The PIV camera is oriented perpendicular to the light sheet, and the processed PIV data yield the two-component velocity field in the plane of the light sheet. In the standard DGV technique, an injection-seeded Nd:YAG pulsed laser light sheet illuminates the seeded flow field, and three receiver systems are used to measure three components of velocity. The receiver systems are oriented at oblique angles to the light sheet in order to accurately resolve the three-component velocity. Each DGV receiver system contains two cameras, which share a common view of the illuminated flow through a beam-splitting cube. One camera views the illuminated flow directly (reference camera) and the second camera images the illuminated flow through an iodine vapor cell (signal camera). The laser frequency (wavelength) is adjusted so that the Doppler-shifted light from particles in the flow falls on an iodine absorption feature, see the following graph. The iodine vapor cell acts as a frequency-to-velocity filter by modulating the intensity of the transmitted light as a function of the flow velocity (Doppler shift). The ratio of the signal and reference images yields the component of the flow velocity along the bisector of the laser sheet propagation direction and the receiver system observation direction. The hybrid system employs a single-component DGV receiver system configured to simultaneously acquire PIV image data, as shown in the following diagram. The cameras used in the DGV receiver are replaced with PIV frame-straddling cameras, and the receiver system views the illuminated light sheet plane at 90 (as in the standard PIV configuration).

Wernet, Mark P.↗

Graphical User Interface for Simulink Integrated Performance Analysis Model

The J-2X Engine (built by Pratt & Whitney Rocketdyne,) in the Upper Stage of the Ares I Crew Launch Vehicle, will only start within a certain range of temperature and pressure for Liquid Hydrogen and Liquid Oxygen propellants. The purpose of the Simulink Integrated Performance Analysis Model is to verify that in all reasonable conditions the temperature and pressure of the propellants are within the required J-2X engine start boxes. In order to run the simulation, test variables must be entered at all reasonable values of parameters such as heat leak and mass flow rate. To make this testing process as efficient as possible in order to save the maximum amount of time and money, and to show that the J-2X engine will start when it is required to do so, a graphical user interface (GUI) was created to allow the input of values to be used as parameters in the Simulink Model, without opening or altering the contents of the model. The GUI must allow for test data to come from Microsoft Excel files, allow those values to be edited before testing, place those values into the Simulink Model, and get the output from the Simulink Model. The GUI was built using MATLAB, and will run the Simulink simulation when the Simulate option is activated. After running the simulation, the GUI will construct a new Microsoft Excel file, as well as a MATLAB matrix file, using the output values for each test of the simulation so that they may graphed and compared to other values.

Durham, R. Caitlyn↗

LSKnowledge: Nexus for Transformative Scientific Discoveries and Enhanced Information Retrieval in NASA Life Sciences Portal

We stand at the brink of an extraordinary transformation in the field of AI, driven by the convergence of generative AI and semantic technologies (e.g., knowledge graphs). This fusion holds immense potential and could redefine the future of scientific exploration, particularly in the realm of life sciences research. In this context, we shed light on the pivotal roles that Large Language Models (LLMs) and semantic technologies will play in advancing research, unearthing and comprehending life sciences information through innovative approaches, and empowering researchers to extract insights from NASA's extensive Life Sciences Data Archive. Within the NASA Life Sciences Portal (NLSP), the integration of LLMs and semantic technologies unlocks several advanced capabilities. First and foremost, it equips scientists with sophisticated tools to manage the ever-expanding wealth of scientific literature and data. Furthermore, it facilitates the creation of knowledge graphs that visually represent intricate relationships among biological entities, enabling comprehensive systems-level analysis. Additionally, the fusion of generative AI (including LLMs) and semantic technology can significantly benefit NASA's life sciences research by enhancing information retrieval and hypothesis generation. These tools enhance natural language understanding, facilitating knowledge discovery within NLSP. The overarching vision is to establish a cohesive knowledge ecosystem within NLSP, harnessing the power of LLMs and semantic technologies to synthesize and cross-reference data from diverse missions, disciplines, and research domains. This holistic approach ultimately deepens our understanding of how space environments impact life sciences data. To advance this initiative, we have launched LSKnowledge, aimed at enhancing the information retrieval capabilities of NLSP. In the short term, our primary goal is to develop a robust semantic search system. This system will empower HRP (Human Research Program) researchers to navigate NLSP data repositories more efficiently and precisely, catalyzing the process of hypothesis formation and scientific breakthroughs. To achieve this, we have employed pre-trained LLMs as part of a semantic search tool that can rank and highlight the most relevant records for user queries. To assess the tool's performance, we have curated a set of approximately 200 queries from subject matter experts (SMEs) and manually ranked the top records retrieved by both the current search system and the new semantic search, using SME judgments as the gold standard for relevancy. Herein, we present the results of our comparative analysis and illustrate how these findings have informed the fine-tuning of the system for enhanced performance. In the long term, our objectives include 1) retrieving publicly available information and integrating it with NLSP data to provide more precise answers to user queries, and 2) incorporating non-textual information from the NLSP database into our approach. In conclusion, the fusion of LLMs and semantic technologies within NLSP represents a pioneering stride towards reshaping the landscape of scientific discovery. This synergy not only equips researchers with powerful tools to navigate the burgeoning sea of information but also facilitates a deeper understanding of complex biological relationships, all while accelerating hypothesis generation and knowledge discovery. Through our initiative, LSKnowledge, we are committed to continually refining and expanding these capabilities, with the aim of not only enhancing information retrieval but also integrating diverse data sources to provide more precise insights. In the grand vision, NLSP strives to become the cornerstone of a comprehensive knowledge ecosystem, unraveling the enigmatic intricacies of life sciences phenomena in the context of space environments.

Life Sciences↗

The Process of the Development of an Operator

On the job training is where new employees called operators can start gaining knowledge on what they will be working on during their time in JSC. In these lessons I learned different things that are important ranging from thermal systems to electrical systems. While doing OJT classes the student will learn how to use a portable computer system which has displays that I also helped edit and clean up. The way you can also learn is by reading system briefs which describes the different systems. Due to the fact of a possible change in the ISS I updated a systems brief so that it can be relevant to what is actually on the space station. I was given a task that will help develop my skills and make myself better prepared for my future in the work field. The project that I worked on had me pulling real time data from the International Space Station. The Data I obtained from the space station will be correlated to battery performance. The group I will be working which is called REBA and we will take the telemetry and evaluate the data. I will be working with my mentor Ben Chislom and co-op Tyler along with the Pro team. They then will put this data into a graph so that they can get the discrepancies and find a way to improve the battery performance. The first weeks I read familiarization books that informed me how the ISS works, how it was built, and the systems that are used to keep the station working. This project is going to benefit NASA by finding out how electricity is being used on the ISS and enabling us to see how it can be used more efficiently. This way we can operate the ISS without wasting power. While conducting research that goes on inside the space station knowing all electricity is being used efficiently.

Banks, Terrence↗

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship. 1. Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. 2. National Academies of Sciences, E. and Medicine, Open Science by Design: Realizing a Vision for 21st Century Research. 2018, Washington, DC: The National Academies Press. 232. 3. Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5.

Life Sciences data↗

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship.

Life Sciences data↗

Identifying Molecules as Biosignatures with Assembly Theory and Mass Spectrometry

The search for evidence of life elsewhere in the universe is hard because it is not obvious what signatures are unique to life. Here we postulate that complex molecules found in high abundance are universal biosignatures as they cannot form by chance. To explore this, we developed the first intrinsic measure of molecular complexity that can be experimentally determined, and this is based upon a new approach called assembly theory which gives the molecular assembly number (MA) of a given molecule. MA allows us to compare the intrinsic complexity of molecules using the minimum number of steps required to construct the molecular graph starting from basic objects, and a probabilistic model shows how the probability of any given molecule forming randomly drops dramatically as its MA increases. To map chemical space, we calculated the MA of ca. 2.5 million compounds, and collected data which showed the complexity of a molecule can be experimentally determined by using three independent techniques including infra-red spectroscopy, nuclear magnetic resonance, and by fragmentation in a mass spectrometer, and this data has an excellent corelation with the values predicted from our assembly theory. We then set out to see if this approach could allow us to identify molecular biosignatures with a set of diverse samples from around the world, outer space, and the laboratory including prebiotic soups. The results show that there is a non-living to living threshold in MA complexity and the higher the MA for a given molecule, the more likely that it had to be produced by a biological process. This work demonstrates it is possible to use this approach to build a life detection instrument that could be deployed on missions to extra-terrestrial locations to detect biosignatures, map the extent of life on Earth, and be used as a molecular complexity scale to quantify the constraints needed to direct prebiotically plausible processes in the laboratory. Such an approach is vital if we are going to find new life elsewhere in the universe or create de-novo life in the lab.

Stuart M Marshall↗

The flow of a compressible fluid past a circular arc profile

The Ackeret iteration process is utilized to obtain higher approximations than that of Prandtl and Glauert for the flow of a compressible fluid past a circular arc profile. The procedure is to expand the velocity potential in a power series of the camber coefficient. The first two terms of the development correspond to the Prandtl-Glauert approximation and yield the well-known correction to the circulation about the profile. The second approximation, involving the square of the camber coefficient, improves the velocity and pressure fields but yields no new results with regard to the circulation, since the circulation about the profile is an odd function of the camber coefficient. The third approximation, involving the cube of the camber coefficient, permits the use of higher values of the camber coefficient and furthermore yields an improvement to the Prandtl-Glauert rule with regard to the effect of compressibility on the circulation of the circular arc profile. Numerical examples with tables and graphs illustrate the results of the analysis.

Kaplan, Carl↗

Telemetry and Science Data Software System

The Telemetry and Science Data Software System (TSDSS) was designed to validate the operational health of a spacecraft, ease test verification, assist in debugging system anomalies, and provide trending data and advanced science analysis. In doing so, the system parses, processes, and organizes raw data from the Aquarius instrument both on the ground and while in space. In addition, it provides a user-friendly telemetry viewer, and an instant pushbutton test report generator. Existing ground data systems can parse and provide simple data processing, but have limitations in advanced science analysis and instant report generation. The TSDSS functions as an offline data analysis system during I&T (integration and test) and mission operations phases. After raw data are downloaded from an instrument, TSDSS ingests the data files, parses, converts telemetry to engineering units, and applies advanced algorithms to produce science level 0, 1, and 2 data products. Meanwhile, it automatically schedules upload of the raw data to a remote server and archives all intermediate and final values in a MySQL database in time order. All data saved in the system can be straightforwardly retrieved, exported, and migrated. Using TSDSS s interactive data visualization tool, a user can conveniently choose any combination and mathematical computation of interesting telemetry points from a large range of time periods (life cycle of mission ground data and mission operations testing), and display a graphical and statistical view of the data. With this graphical user interface (GUI), the data queried graphs can be exported and saved in multiple formats. This GUI is especially useful in trending data analysis, debugging anomalies, and advanced data analysis. At the request of the user, mission-specific instrument performance assessment reports can be generated with a simple click of a button on the GUI. From instrument level to observatory level, the TSDSS has been operating supporting functional and performance tests and refining system calibration algorithms and coefficients, in sync with the Aquarius/SAC-D spacecraft. At the time of this reporting, it was prepared and set up to perform anomaly investigation for mission operations preceding the Aquarius/SAC-D spacecraft launch on June 10, 2011.

Bates, Lakesha↗

Surrogate oracles, generalized dependency and simpler models

Software reliability models require the sequence of interfailure times from the debugging process as input. It was previously illustrated that using data from replicated debugging could greatly improve reliability predictions. However, inexpensive replication of the debugging process requires the existence of a cheap, fast error detector. Laboratory experiments can be designed around a gold version which is used as an oracle or around an n-version error detector. Unfortunately, software developers can not be expected to have an oracle or to bear the expense of n-versions. A generic technique is being investigated for approximating replicated data by using the partially debugged software as a difference detector. It is believed that the failure rate of each fault has significant dependence on the presence or absence of other faults. Thus, in order to discuss a failure rate for a known fault, the presence or absence of each of the other known faults needs to be specified. Also, in simpler models which use shorter input sequences without sacrificing accuracy are of interest. In fact, a possible gain in performance is conjectured. To investigate these propositions, NASA computers running LIC (RTI) versions are used to generate data. This data will be used to label the debugging graph associated with each version. These labeled graphs will be used to test the utility of a surrogate oracle, to analyze the dependent nature of fault failure rates and to explore the feasibility of reliability models which use the data of only the most recent failures.

Wilson, Larry↗