Engineering PapersSearch

SEARCH · Engineering Papers

Results for “ML”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY

Applying Transfer Learning for Street-Scale Nuisance Flood Forecasting in Coastal-Urban Cities

An important challenge with Machine Learning (ML) is its transferability; that is, whether a ML model trained on one set of data can be applied to a second set of data without requiring a full re-training of the model. Transfer Learning (TL) addresses this challenge by transferring knowledge learned in the source domain (the data it was trained on) to the target domain (a second set of data that is statistically different but related, which the model was not trained on). This study investigates the use of TL for street-scale nuisance flood forecasting by exploring whether a ML model trained on data collected for one set of streets can effectively forecast flooding for another set of streets in the same city using TL. The envisioned use case is a city deploying a new flood depth monitoring sensor on a street and using TL to apply a ML model, trained on sensor data from an existing flood depth sensor network, to this new street. Eventually, the new flood depth sensor will have a sufficient dataset for training its own ML model, but TL can be used to fill the gap in time while this new dataset is being generated. This method is explored using a Long Short-Term Memory (LSTM) model trained on data for the flood-prone streets of Norfolk City, Virginia. The data used for training includes environmental time series (rainfall, tide), topographic features (Digital Elevation Model (DEM), Topographic Wetness Index (TWI), Depth To Water (DTW)), and street-scale flood depth time series obtained from a high-fidelity physics-based model, acting as a synthetic street-scale stream depth sensor dataset since actual stream depth sensor data is generally unavailable for most cities. A set of 180 flood-prone streets was used to train a base model, while another set of 180 flood-prone streets was used to re-train that model using different TL strategies. The results show that full-weight re-training proved most effective and minimal re-training of only the output layer was insufficient. The advantage of TL was most pronounced when target data was limited, meaning data collected at the new water depth sensor location included generally less than 18 flood events. As target data increased beyond 18 flood events, the benefit of TL diminished relative to training a ML model directly on the local flood events. These findings can assist cities as they implement street-scale flood sensing systems to create accurate forecasts for new sensing locations that do not yet have sufficient data records to train a local ML model.

Roy, Binata [Univ. of Virginia, Charlottesville, V

Best estimate of the planetary boundary layer height from multiple remote sensing measurements

Remote sensing measurements have been widely used to estimate the planetary boundary layer height (PBLHT). Each remote sensing approach offers unique strengths and faces different limitations. In this study, we use machine learning (ML) methods to produce a best-estimate PBLHT (PBLHT-BE-ML) by integrating four PBLHT estimates derived from remote sensing measurements at the Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) observatory. Three ML models – random forest (RF) classifier, RF regressor, and light gradient-boosting machine (LightGBM) – were trained on a dataset from 2017 to 2023 that included radiosonde, various remote sensing PBLHT estimates, and atmospheric meteorological conditions. Evaluations indicated that PBLHT-BE-ML from all three models improved alignment with the PBLHT derived from radiosonde data (PBLHT-SONDE), with LightGBM demonstrating the highest accuracy under both stable and unstable boundary layer conditions. Feature analysis revealed that the most influential input features at the SGP site were the PBLHT estimates derived from (a) potential temperature profiles retrieved using Raman lidar (RL) and atmospheric emitted radiance interferometer (AERI) measurements (PBLHT-THERMO), (b) vertical velocity variance profiles from Doppler lidar (PBLHT-DL), and (c) aerosol backscatter profiles from micropulse lidar (PBLHT-MPL). The trained models were then used to predict PBLHT-BE-ML at a temporal resolution of 10 min, effectively capturing the diurnal evolution of PBLHT and its significant seasonal variations, with the largest diurnal variation observed over summer at the SGP site. We applied these trained models to data from the ARM Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) field campaign (EPC), where the PBLHT-BE-ML, particularly with the LightGBM model, demonstrated improved accuracy against PBLHT-SONDE. Analyses of model performance at both the SGP and EPC sites suggest that expanding the training dataset to include various surface types, such as ocean and ice-covered areas, could further enhance ML model performance for PBLHT estimation across varied geographic regions.

Zhang, Damao [Pacific Northwest National Laborator

Automated Desalting Apparatus

Because salt and metals can mask the signature of a variety of organic molecules (like amino acids) in any given sample, an automated system to purify complex field samples has been created for the analytical techniques of electrospray ionization/ mass spectroscopy (ESI/MS), capillary electrophoresis (CE), and biological assays where unique identification requires at least some processing of complex samples. This development allows for automated sample preparation in the laboratory and analysis of complex samples in the field with multiple types of analytical instruments. Rather than using tedious, exacting protocols for desalting samples by hand, this innovation, called the Automated Sample Processing System (ASPS), takes analytes that have been extracted through high-temperature solvent extraction and introduces them into the desalting column. After 20 minutes, the eluent is produced. This clear liquid can then be directly analyzed by the techniques listed above. The current apparatus including the computer and power supplies is sturdy, has an approximate mass of 10 kg, and a volume of about 20 20 20 cm, and is undergoing further miniaturization. This system currently targets amino acids. For these molecules, a slurry of 1 g cation exchange resin in deionized water is packed into a column of the apparatus. Initial generation of the resin is done by flowing sequentially 2.3 bed volumes of 2N NaOH and 2N HCl (1 mL each) to rinse the resin, followed by .5 mL of deionized water. This makes the pH of the resin near neutral, and eliminates cross sample contamination. Afterward, 2.3 mL of extracted sample is then loaded into the column onto the top of the resin bed. Because the column is packed tightly, the sample can be applied without disturbing the resin bed. This is a vital step needed to ensure that the analytes adhere to the resin. After the sample is drained, oxalic acid (1 mL, pH 1.6-1.8, adjusted with NH4OH) is pumped into the column. Oxalic acid works as a chelating reagent to bring out metal ions, such as calcium and iron, which would otherwise interfere with amino acid analysis. After oxalic acid, 1 mL 0.01 N HCl and 1 mL deionized water is used to sequentially rinse the resin. Finally, the amino acids attached to the resin, and the analytes are eluted using 2.5 M NH4OH (1 mL), and the NH4OH eluent is collected in a vial for analysis.

Spencer, Maegan K.

A Novel Protocol for Decoating and Permeabilizing Bacterial Spores for Epifluorescent Microscopy

Based on previously reported procedures for permeabilizing vegetative bacterial cells, and numerous trial-and-error attempts with bacterial endospores, a protocol was developed for effectively permeabilizing bacterial spores, which facilitated the applicability of fluorescent in situ hybridization (FISH) microscopy. Bacterial endospores were first purified from overgrown, sporulated suspensions of B. pumilus SAFR-032. Purified spores at a concentration of approx equals 10 million spores/mL then underwent proteinase-K treatment, in a solution of 468.5 μL of 100 mM Tris-HCl, 30 μL of 10% SDS, and 1.5 microL of 20 mg/mL proteinase-K for ten minutes at 35 ºC. Spores were then harvested by centrifugation (15,000 g for 15 minutes) and washed twice with sterile phosphate-buffered saline (PBS) solution. This washing process consisted of resuspending the spore pellets in 0.5 mL of PBS, vortexing momentarily, and harvesting again by centrifugation. Treated and washed spore pellets were then resuspended in 0.5 mL of decoating solution, which consisted of 4.8 g urea, 3 mL Milli-Q water, 1 mL 0.5M Tris, 1 mL 1M dithiothreitol (DTT), and 2 mL 10% sodium-dodecylsulfate (SDS), and were incubated at 65 ºC for 15 minutes while being shaken at 165 rpm. Decoated spores were then, once again, washed twice with sterile PBS, and subjected to lysozyme/mutanolysin treatment (7 mg/mL lysozyme and 7U mutanolysin) for 15 minutes at 35 C. Spores were again washed twice with sterile PBS, and spore pellets were resuspended in 1-mL of 2% SDS. This treatment, facilitating inner membrane permeabilization, lasted for ten minutes at room temperature. Permeabilized spores were washed two final times with PBS, and were resuspended in 200 mkcroL of sterile PBS. At this point, the spores were permeable and ready for downstream processing, such as oligonucleotideprobe infiltration, hybridization, and microscopic evaluation. FISH-microscopic imagery confirmed the effective and efficient (≈50% successful permeabilization and recovery) permeabilization of numerous spore preparations. The novelty of the technology developed here is in its applicability to bacterial endospores. While protocols abound for the effective permeabilization of bacterial, archaeal, and eukaryotic vegetative cells, there are no such reliable methods for decoating and permeabilizing bacterial endospores in a manner that is amenable to downstream FISH microscopic analyses. This innovation enables the direct visualization and enumeration of spores via FISH-based microscopic techniques, circumventing the complications that accompany previously required germination regimes. The synergistic enzymatic weakening of the many spore layers facilitates a structural compromise that is just enough to render the spores permeable without degrading the spore to a level, which precludes it from recognition.

LaDuc, Myron T.

Modal Test of the NASA Mobile Launcher at Kennedy Space Center

The NASA Mobile Launcher (ML), located at Kennedy Space Center (KSC), has recently been modified to support the launch of the new NASA Space Launch System (SLS). The ML is a massive structure—consisting of a 345-foot tall tower attached to a two-story base, weighing approximately 10.5 million pounds—that will secure the SLS vehicle as it rolls to the launch pad on a Crawler Transporter, as well as provide a launch platform at the pad. The ML will also provide the boundary condition for an upcoming SLS Integrated Modal Test (IMT). To help correlate the ML math models prior to this modal test, and allow focus to remain on updating SLS vehicle models during the IMT, a ML-only experimental modal test was performed in June 2019. Excitation of the tower and platform was provided by five uniquely-designed test fixtures, each enclosing a hydraulic shaker, capable of exerting thousands of pounds of force into the structure. For modes not that were not sufficiently excited by the test fixture shakers, a specially-designed mobile drop tower provided impact excitation at additional locations of interest. The response of the ML was measured with a total of 361 accelerometers. Following the random vibration, sine sweep vibration, and modal impact testing, frequency response functions were calculated and modes were extracted for three different configurations of the ML in 0 Hz to 12 Hz frequency range. This paper will provide a case study in performing modal tests on large structures by discussing the Mobile Launcher, the test strategy, an overview of the test results, and recommendations for meeting a tight test schedule for a large-scale modal test.

Hydraulic shakers

Fusion of Test and Analysis: Artemis I Booster to Mobile Launcher Interface Validation

NASA is in the midst of bold and exciting next steps in human exploration and spaceflight. The designs of the new Space Launch System (SLS), the Orion spacecraft and the Exploration Ground Systems (EGS) for vehicle processing and launch are essentially complete and there has been significant progress in manufacturing and assembly of specific hardware for the Artemis I and Artemis II missions. Equally as important, the program level and integrated system level testing and analyses are also well underway to support integrated verification, validation, and Certificate of Flight Readiness (CoFR) for Artemis I. Testing and analysis are key to addressing technical challenges faced by the Artemis missions. Building block approaches are required that provide the right balance between component, element, and/or system level testing that satisfies verification and validation objectives where uncertainties are quantified and minimized. Artemis I is a system of systems that requires a fusion of test and analysis that adeptly characterizes critical interfaces between major program elements. An example of this fusion involves characterizing the interface between the SLS booster and the Mobile Launcher (ML) Vertical Support Post (VSP) interfaces. Proper characterization of this interface represents a number of challenges beginning with the fact that it is a mating of ground support structure in the form of a civil structure to flight hardware. Both sides of the interface are built to different construction standards, but are governed by interface requirements to ensure compatibility when mated. From past program experience, the flexibility at the booster to ML interface is critical in developing accurate prelaunch stacking and cryogenic preloads, squat loads, and pad separation release of preloads and squat loads. This same premise holds for Artemis I. To characterize the asymmetric characteristics at this interface, careful consideration of static forces due to gravity loading with the commensurate effects due to leveling during booster stacking (i.e., spacing and shimming) and nonlinear geometric forces are necessary for inclusion in pre-test assessments. This paper will look at these issues for the upcoming Booster Pull Test in which two boosters will be installed on the ML and one of these boosters will undergo static lateral loading followed afterwards with dynamic excitation into resonance and free-decay. This paper evaluates the booster to ML interface characteristics by characterizing the interface flexibility between the booster aft skirt and the ML VSP interfaces. Furthermore, this paper methodically evaluates the effect of the following on the test outcome: gravitational effects on the booster and ML, the effects of VSP leveling, spacing, and shimming under gravitational loading during booster stacking, the effect of geometric nonlinear follower force due to cg offset as booster is laterally displaced, and the system coupling between the booster under test, ML, and the second booster. Simulated results for a static load pull and dynamic excitation provide insight into the differences in measurement responses when boundary conditions and geometric conditions are included and not included.

Joel W Sills Jr.

Toward Certification of Machine-Learning Systems for Low Criticality Airborne Applications

The exceptional progress in the field of machine learning (ML) in recent years has attracted a lot of interest in using this technology in aviation. Possible airborne applications of ML include safety-critical functions, which must be developed in compliance with rigorous certification standards of the aviation industry. Current certification standards for the aviation industry were developed prior to the ML renaissance without taking specifics of ML technology into account. There are some fundamental incompatibilities between traditional design assurance approaches and certain aspects of ML-based systems. In this paper, we analyze the current airborne certification standards and show that all objectives of the standards can be achieved for a low-criticality ML-based system if certain assumptions about ML development workflow are applied.

Avionics

Machine-Learning for Safety Critical Airborne Applications Part II: Case Study

The exceptional progress in the field of Artificial Intelligence (AI) systems, enabled by Machine Learning (ML) technology in recent years provides historic opportunities for the aviation industry. Current certification standards for avionics were developed prior to the ML renaissance and have several fundamental incompatibilities with the ML technology. WG-114 is working hard to release a new standard as soon as possible but for now there is no recognized means of compliance for ML based systems even of low criticality. In this talk, we present the custom ML workflow that can be used comply with all objectives of the current certification standards for a low-criticality (DAL D and C) ML-based system. To illustrate the practical application of the custom ML workflow we present a case study of a system based on a Deep Neural Network (DNN) intended to detect and identify airport runway signs. We present the system design, data generation, training, and verification in detail and describe how the design assurance objectives can be met for a DAL D and DAL C systems.

Johann Schumann

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning

Intracranial Effects of Intermittent Lower Body Negative Pressure with Head Down Tilt Bed Rest: Comparison to Upright Posture

INTRODUCTION Spaceflight-associated neuro-ocular syndrome (SANS) affects a majority of astronauts during long-duration spaceflight. SANS is hypothesized to result from an unrelenting headward fluid shift that occurs in the microgravity environment. Altered intracranial structure and physiology documented in spaceflight and head-down tilt (HDT) experiments(1, 2) are also theorized to be related to chronic headward fluid shift and thereby can provide an independent quantitative assessment related to this mechanism. As a potential countermeasure, lower body negative pressure (LBNP) applied in the supine posture has shown efficacy in reducing headward fluid shift(3). However, it remains unknown if intermittent LBNP simulating daily upright posture fluid redistribution can mitigate SANS and intracranial changes associated with long-duration spaceflight. The goal of this study was to quantify the effects of daily application of LBNP or daily exposure to the upright posture on intracranial structure and physiology during long-term HDT in order to evaluate the potential efficacy of LBNP as a countermeasure to SANS. METHODS 22 healthy volunteers (11 men; 11 women; mean age = 35 years (SD=9.2)(age range = 24 to 52 years); mean BMI = 24.0 kg/m2 (SD=2.8)) completed an MRI study performed at the German Aerospace Facility in Cologne, Germany. Strict six-degree head-down tilt (HDT) bedrest was used as a spaceflight analog to induce a continuous headward fluid shift for 30 days. The subjects were divided equally into two groups of interventions: 1) LBNP for 3-hour sessions twice daily and 2) seated position for 3-hour sessions twice daily. Interventions were divided into morning and afternoon sessions for both groups. The LBNP intervention was maintained at 25 mmHg  2 mmHg. Pulse-gated MRI phase-contrast flow imaging was used to quantify cerebral artery stroke volume (CASV) and peak-to-peak cerebral spinal fluid (CSF) velocity (CSFVp-p) within the cerebral aqueduct. A 3D T1-MPRAGE sequence was used to quantify volumetric changes of the brain and intracranial CSF spaces using MRI Cloud software. MRI acquisitions were obtained at baseline (BDC, supine posture), 15 days into HDT (HDT15), 29 days into HDT (HDT29), and 12 days after recovery (R12, supine posture). The data were analyzed by a mixed model, which included intervention, time (four nominal levels BDC, HDT15, HDT29, R12), and intervention-time interaction as the fixed effects and included subject as a random effect. RESULTS Compared to BDC there was no statistically significant difference in CASV and CSFVp-p during HDT except for CASV at HDT29 (seated group) where there was a 1.7 mL (12%) decrease (P<0.01). Compared to BDC there was a statistically significant increase in intracranial volume (ICV; ICV = white matter + gray matter + CSF) for both interventions at HDT15 (Δ13 mL, 0.9 % (LBNP group); Δ15 mL, 1.0%, (seated group)) and HDT29 (Δ19 mL, 1.2%, (LBNP group); Δ23 mL, 1.5%) (seated group)) (All Ps<.001). Compared to baseline, lateral ventricular volume increased at HDT15 (Δ0.9 mL, 5.5%) and HDT29 (Δ1.8 mL, 10%)) for LBNP only (P=.001 for each). During HDT, white matter volume remained stable compared to baseline for both interventions. There were no significant intervention effects in the overall response to HDT (All Ps >.2). CONCLUSION LBNP had a similar response to the seated posture during 29 days of HDT. Although there was an increase in ICV and lateral ventricular volume with HDT there was no change in intracranial physiological parameters with LBNP suggesting an overall diminished response to the long-term effects of HDT. The lack of any significant increase in white matter volume during HDT with LBNP suggests maintenance of cerebral interstitial fluid transport. Final conclusions are pending the data collection for the control group (no intervention) which will be completed in 2023.

Larry A Kramer

A Robust Machine Learning Schema for Developing, Maintaining, and Disseminating Machine Learning Models

Recent advances in the development of machine learning (ML) algorithms have enabled the creation of predictive models that can improve decision making, decrease computational cost, and improve efficiency in a variety of fields. As an organization begins to develop and implement such models, the data used in the training, validation, and testing of ML models, the model parameters, and the use cases or limitations of the models must be properly stored to ensure models are both fully traceable and used correctly. In the context of predicting material behavior, advances in computationally intense, physics-based modeling of material behavior at various length scales and the emergence of Integrated Computational Materials Engineering (ICME) have driven the need for developing data-driven surrogate models of the physics-based simulation tools using ML techniques. Surrogate model development allows for accurate material behavior prediction at a fraction of the cost of its physics-based counterpart, allowing for multiscale simulations of real-world applications, further enabling the ability to design fit-for-purpose materials for a reasonable computational investment. However, training such models requires extensive data, and thus, effective data management is necessary to reach the full potential that ML can offer to material design and ICME. This paper proposes a generalized, robust schema that allows organizations to store both real (experimental) and virtual (simulation) data used to train ML models and the defining model parameters and architectures within the Granta MI Platform. The developed schema allows for various types of data inputs and outputs, including single point values, time-series data, and images that can be used in the prediction of material behavior, while following outlined best practices for effective data management. An effective schema for ML data and models can help prevent the recreation of virtual/real training data and surrogate models, help reduce the time to create new models similar to existing ones by offering a starting point in the hyperparameter determination stages, minimize resources devoted to verification and validation (V&V) and certification of models, and ensure that data and surrogate models are not misused due to full traceability of both the data and ML model. It also allows organizations access to models that have already been developed, such that they can be used in the design of new materials, enabling the overall goals of ICME.

Brandon L. Hearley

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of machine-learning (ML) models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a user-friendly app that has been created to deploy trained machine-learning (ML) models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of ML application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). The app gets user input for a turbofan design, preprocesses the input data, and deploys trained ML models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The ML predictive models were built by employing supervised deep-learning and K-nearest neighbor regression algorithms to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these ML models using the app shows that Aero-Engines AI is an easy-touse and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage. Current version of the app focuses on the performance prediction of conventional turbofans. However, the scope of the app can easily be expanded to include other engine types (such as turboshaft and hybrid-electric systems) after their ML models are developed. Overall, the use of a machine-learning app for aircraft engine concept assessment represents a promising area of development in aircraft engine conceptual design.

machine learning

Aero-Engines AI - A Machine-Learning App for Aircraft Engine Concepts Assessment

Effective deployment of machine-learning (ML) models could drive a high level of efficiency in aircraft engine conceptual design. Aero-Engines AI is a user-friendly app that has been created to deploy trained machine-learning (ML) models to assess aircraft engine concepts. It was created using tkinter, a GUI (graphical user interface) module that is built into the standard Python library. Employing tkinter greatly facilitates the sharing of ML application as an executable file which can be run on Windows machines (without the need to have Python or any library installed). The app gets user input for a turbofan design, preprocesses the input data, and deploys trained ML models to predict turbofan thrust specific fuel consumption (TSFC), engine weight, core size, and turbomachinery stage-counts. The ML predictive models were built by employing supervised deep-learning and K-nearest neighbor regression algorithms to study patterns in an existing open-source database of production and research turbofan engines. They were trained, cross-validated, and tested in Keras, an open-source neural networks API (application programming interface) written in Python, with TensorFlow (Google open-source artificial intelligence library) serving as the backend engine. The smooth deployment of these ML models using the app shows that Aero-Engines AI is an easy-touse and a time-saving tool for aircraft engine design-space exploration during the conceptual design stage. Current version of the app focuses on the performance prediction of conventional turbofans. However, the scope of the app can easily be easily expanded to include other engine types (such as turboshaft and hybrid-electric systems) after their ML models are developed. Overall, the use of a machine-learning app for aircraft engine concept assessment represents a promising area of development in aircraft engine conceptual design.

machine learning

The Development and Deployment of Machine Learning Models for Aircraft Engine Concept Assessment

In today's competitive landscape, the effective development and utilization of machine-learning (ML) applications have become imperative across various sectors. This study presents an outline of the procedure involved in creating and implementing ML models for conceptualizing and evaluating aircraft engines. These models leverage supervised deep-learning algorithms to analyze patterns within an open-source repository containing data on both production and research conventional turbofan engines. The main areas of focus encompass crucial engine parameters like thrust-specific fuel consumption (TSFC), engine weight, engine diameter, and turbomachinery stage counts. While the creation of ML models is fundamental for their utilization, ensuring their seamless deployment holds equal significance. To address this aspect, a conversational AI chatbot is constructed, utilizing natural language processing (NLP) techniques, to facilitate the deployment of these ML models. The comprehensive workflow encompasses several key stages: gathering and enhancing engine data, training and cross validating the ML models, testing and evaluating their performance, and finally, deploying, monitoring, and updating the ML models. By following this systematic approach, the aim is to streamline the development and deployment process of ML models tailored for aircraft engine assessment.

Development

The Development and Deployment of Machine Learning Models for Aircraft Engine Concept Assessment

In today's competitive landscape, the effective development and utilization of machine-learning (ML) applications have become imperative across various sectors. This study presents an outline of the procedure involved in creating and implementing ML models for conceptualizing and evaluating aircraft engines. These models leverage supervised deep-learning algorithms to analyze patterns within an open-source repository containing data on both production and research conventional turbofan engines. The main areas of focus encompass crucial engine parameters like thrust-specific fuel consumption (TSFC), engine weight, engine diameter, and turbomachinery stage counts. While the creation of ML models is fundamental for their utilization, ensuring their seamless deployment holds equal significance. To address this aspect, a conversational AI chatbot is constructed, utilizing natural language processing (NLP) techniques, to facilitate the deployment of these ML models. The comprehensive workflow encompasses several key stages: gathering and enhancing engine data, training and cross validating the ML models, testing and evaluating their performance, and finally, deploying, monitoring, and updating the ML models. By following this systematic approach, the aim is to streamline the development and deployment process of ML models tailored for aircraft engine assessment.

Development