Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mixture modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Scalable probabilistic estimates of electric vehicle charging given observed driver behavior

To prepare for rapid growth in global electric vehicle adoption, grid and policy planners depend on detailed forecasts of future charging demand. In this paper we propose a novel holistic, scalable, probabilistic framework to produce large-scale estimates of electric vehicle charging load for long-term planning that capture real drivers’ charging patterns. Our framework captures the uncertainty and stochasticity in charging demand by taking a graphical modeling approach. It has three core elements: driver groups, charging segment choices, and charging session time and energy requirements. The framework uses hierarchical clustering to group drivers by their charging histories, capturing their heterogeneous behaviors and preferences across different segments or types of charging. The framework uses probabilistic mixture models for each driver group’s sessions to identify the unique charging behaviors observed within each segment. We illustrate its application with a large data set from California, profiling the charging patterns and unique driver clusters it identifies. Using the model knobs representing drivers’ battery capacities, behavior, and segment access we present scenarios for California’s charging demand in 2030 with 8 million passenger electric vehicles. Peak charging demand ranged from 3.3 to 8.7 GW across scenarios. Furthermore, each was calculated in under 45 s on a laptop computer.

33 ADVANCED PROPULSION SYSTEMS↗

Automatic Generation of Algorithms for the Statistical Analysis of Planetary Nebulae Images

Analyzing data sets collected in experiments or by observations is a Core scientific activity. Typically, experimentd and observational data are &aught with uncertainty, and the analysis is based on a statistical model of the conjectured underlying processes, The large data volumes collected by modern instruments make computer support indispensible for this. Consequently, scientists spend significant amounts of their time with the development and refinement of the data analysis programs. AutoBayes [GF+02, FS03] is a fully automatic synthesis system for generating statistical data analysis programs. Externally, it looks like a compiler: it takes an abstract problem specification and translates it into executable code. Its input is a concise description of a data analysis problem in the form of a statistical model as shown in Figure 1; its output is optimized and fully documented C/C++ code which can be linked dynamically into the Matlab and Octave environments. Internally, however, it is quite different: AutoBayes derives a customized algorithm implementing the given model using a schema-based process, and then further refines and optimizes the algorithm into code. A schema is a parameterized code template with associated semantic constraints which define and restrict the template s applicability. The schema parameters are instantiated in a problem-specific way during synthesis as AutoBayes checks the constraints against the original model or, recursively, against emerging sub-problems. AutoBayes schema library contains problem decomposition operators (which are justified by theorems in a formal logic in the domain of Bayesian networks) as well as machine learning algorithms (e.g., EM, k-Means) and nu- meric optimization methods (e.g., Nelder-Mead simplex, conjugate gradient). AutoBayes augments this schema-based approach by symbolic computation to derive closed-form solutions whenever possible. This is a major advantage over other statistical data analysis systems which use numerical approximations even in cases where closed-form solutions exist. AutoBayes is implemented in Prolog and comprises approximately 75.000 lines of code. In this paper, we take one typical scientific data analysis problem-analyzing planetary nebulae images taken by the Hubble Space Telescope-and show how AutoBayes can be used to automate the implementation of the necessary anal- ysis programs. We initially follow the analysis described by Knuth and Hajian [KHO2] and use AutoBayes to derive code for the published models. We show the details of the code derivation process, including the symbolic computations and automatic integration of library procedures, and compare the results of the automatically generated and manually implemented code. We then go beyond the original analysis and use AutoBayes to derive code for a simple image segmentation procedure based on a mixture model which can be used to automate a manual preproceesing step. Finally, we combine the original approach with the simple segmentation which yields a more detailed analysis. This also demonstrates that AutoBayes makes it easy to combine different aspects of data analysis.

Fischer, Bernd↗

Probabilistic Solar Power Forecasting Using Bayesian Model Averaging

There is rising interest in probabilistic forecasting to mitigate risks from solar power uncertainty, but the numerical weather prediction (NWP) ensembles readily available to system operators are often biased and underdispersed. We propose a Bayesian model averaging (BMA) post-processing method suitable for forecasting power from utility-scale photovoltaic (PV) plants at multiple time horizons up to at least the day-ahead timescale. BMA is a kernel dressing technique for NWP ensembles in which the forecast is a weighted sum of member-specific probability density functions. We tailor BMA for utility-scale PV forecasting by modeling power clipping at the AC inverter rating and advance the theory of BMA with a new beta kernel parameterization that accommodates theoretical constraints not previously addressed. BMA is demonstrated for a case study of 11 utility-scale PV plants in Texas, forecasting at hourly resolution for the complete year 2018. BMA's mixture-model approach mitigates underdispersion of the raw ensemble to significantly improve forecast calibration, while consistently outperforming an ensemble model output statistics (EMOS) parametric approach from the literature. At 4-hour lead time, the BMA post-processing achieves continuous ranked probability skill scores of 2--36% over the raw ensemble, with consistent performance at multiple lead times suitable for power system operations.

14 SOLAR ENERGY↗

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors, such as the capabilities and the performance of the accelerator and detectors, the use and exploitation of the available information, the design of search strategies and observables, as well as the proposal of new models. We focus on the use of the information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning techniques and tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by pp\to hh\to b\bar b b \bar b p p → h h → b b ‾ b b ‾ , selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with 1% and 0.5% true signal fractions in the dataset. We also show that the method is robust against the absence of signal. We discuss potential prospects for taking this Bayesian data-driven paradigm into more realistic scenarios.

Alvarez, Ezequiel↗

Modelling changes in the dielectric and scattering properties of young snow-covered sea ice at GHz frequencies

Observations of the physical properties of the snow cover and underlying young fast ice in Resolute Passage, Canada, were made during the winter of 1982. Detailed measurements of snow density and ice and snow temperatures, salinities, and brine volumes were made over a period of 46 d, beginning when the ice was 0.4 m thick and about 9 d old. The recorded values are used in a theoretical mixture model to predict the dielectric properties of the snow cover over the microwave frequency range. The results of this analysis are then used to investigate the effects of the snow properties on the radar backscatter signatures of young sea ice. The results show that backscatter is a function of the incidence angle and can change significantly over short periods of time during the early evolutionary phase of ice and snow-cover development. This has important consequences for the identification of young ice forms from SAR or SLAR images.

Drinkwater, Mark R.↗

Using Neural Networks to Identify Mixture Components in Hyperspectral Reflectance Data

Neural networks have been employed to identify materials of interest from hyperspectral data (generally imagery) based on their unique spectral signatures. This approach assumes that there is a single material that is standing out from the rest of the spectrum to be identified. However, pixels often contain more than one material, or a material of interest may itself be a mixture of multiple materials. Neural networks are only as good as the data used to train them, and it takes a great deal of work in the laboratory to identify, make, and measure all potential mixtures of interest. Thus, researchers often calculate synthetic spectra using algorithms with varying degrees of fidelity to the physics that govern the interactions between light and multiple materials. In this work, we have (1) adapted a neural network designed to identify mixture components from Raman spectroscopy to work with visible to near‐infrared reflectance data and (2) tested three common mixture algorithms to determine the most accurate and least computationally expensive method to build synthetic training datasets. With our initial test dataset, we have achieved accuracies of > 90% and found that the synthetic training dataset produced using the Hapke mixture model provides the best results.

99 GENERAL AND MISCELLANEOUS↗

Molecular dipole moment learning via rotationally equivariant derivative kernels in molecular-orbital-based machine learning

This study extends the accurate and transferable molecular-orbital-based machine learning (MOB-ML) approach to modeling the contribution of electron correlation to dipole moments at the cost of Hartree–Fock computations. A MOB pairwise decomposition of the correlation part of the dipole moment is applied, and these pair dipole moments could be further regressed as a universal function of MOs. The dipole MOB features consist of the energy MOB features and their responses to electric fields. An interpretable and rotationally equivariant derivative kernel for Gaussian process regression (GPR) is introduced to learn the dipole moment more efficiently. The proposed problem setup, feature design, and ML algorithm are shown to provide highly accurate models for both dipole moments and energies on water and 14 small molecules. To demonstrate the ability of MOB-ML to function as generalized density-matrix functionals for molecular dipole moments and energies of organic molecules, we further apply the proposed MOB-ML approach to train and test the molecules from the QM9 dataset. The application of local scalable GPR with Gaussian mixture model unsupervised clustering GPR scales up MOB-ML to a large-data regime while retaining the prediction accuracy. In addition, compared with the literature results, MOB-ML provides the best test mean absolute errors of 4.21 mD and 0.045 kcal/mol for dipole moment and energy models, respectively, when training on 110 000 QM9 molecules. The excellent transferability of the resulting QM9 models is also illustrated by the accurate predictions for four different series of peptides.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Prospects for using carbon-carbon composites for EMI shielding

Since pyrolyzed carbon has a higher electrical conductivity than most polymers, carbon-carbon composites would be expected to have higher electromagnetic interference (EMI) shielding ability than polymeric resin composites. A rule of mixtures model of composite conductivity was used to calculate the effect on EMI shielding of substituting a pyrolyzed carbon matrix for a polymeric matrix. It was found that the improvements were small, no more than about 2 percent for the lowest conductivity fibers (ex-rayon) and less than 0.2 percent for the highest conductivity fibers (vapor grown carbon fibers). The structure of the rule of mixtures is such that the matrix conductivity would only be important in those cases where it is much higher than the fiber conductivity, as in metal matrix composites.

Gaier, James R.↗

Application of Gaussian Mixture Regression for the Correction of Low Cost PM2.5 Monitoring Data in Accra, Ghana

Low-cost sensors (LCSs) for air quality monitoring have enormous potential to improve air quality data coverage in resource-limited parts of the world such as sub-Saharan Africa. LCSs, however, are affected by environment and source conditions. To establish high-quality data, LCSs must be collocated and calibrated with reference grade PM2.5 monitors. From March 2020, a low-cost PurpleAir PM2.5 monitor was collocated with a Met One Beta Attenuation Monitor 1020 in Accra, Ghana. While previous studies have shown that multiple linear regression (MLR) and random forest regression (RF) can improve accuracy and correlation between PurpleAir and reference data, MLR and RF yielded suboptimal improvement in the Accra collocation (R2 = 0.81 and R2 = 0.81, respectively). We present the first application of Gaussian mixture regression (GMR) to air quality data calibration and demonstrate improvement over traditional methods by increasing the collocated PM2.5 correlation and accuracy to R2 = 0.88 and MAE = 2.2 μg/cu. m. Gaussian mixture models (GMMs) are a probability density estimator and clustering method from which nonlinear regressions that tolerate missing inputs can be derived. We find that even when given missing inputs, GMR provides better correlation than MLR and RF performed with complete data. GMR also allows us to estimate calibration certainty. When evaluated, 95% confidence intervals agreed with reference PM2.5 data 96% of the time, suggesting that the model accurately assesses its own confidence. Additionally, clustering within the GMM is consistent with climate characteristics, providing confidence that the calibration approach can learn underlying relationships in data.

Sensors↗

The Role of the Anion in Imidazolium-Based Ionic Liquids for Fuel and Terpenes Processing

The potentialities of methylimidazolium-based ionic liquids (ILs) as solvents were evaluated for some relevant separation problems—terpene fractionation and fuel processing—studying selectivities, capacities, and solvent performance indices. The activity coefficients at infinite dilution of the solute (1) in the IL (3), γ13∞, of 52 organic solutes were measured by inverse gas chromatography over a temperature range of 333.2–453.2 K. The selected ILs are 1-butyl-3-methylimidazolium hexafluorophosphate, [C4mim][PF6], and the equimolar mixture of [C4mim][PF6] and 1-butyl-3-methylimidazolium chloride, [C4mim]Cl. Generally, low polar solutes follow γ1,C4mimCl∞ > γ1,C4mimPF6+C4mimCl∞ > γ1,C4mimPF6∞ while the opposite behavior is observed for alcohols and water. For citrus essential oil deterpenation, the results suggest that cations with long alkyl chains, such as C12mim+, promote capacity, while selectivity depends on the solute polarity. Promising results were obtained for the separation of several model mixtures relevant to fuel industries using the equimolar mixture of [C4mim][PF6] and [C4mim]Cl. This work demonstrates the importance of tailoring the polarity of the solvents, suggesting the use of ILs with mixed anions as alternative solvents for the removal of aliphatic hydrocarbons and contaminants from fuels.

Zambom, Aline (ORCID:0000000324512716)↗

From minerals to rocks: Toward modeling lithologies with remote sensing

High spectral resolution imaging spectroscopy will play an important role in future planetary missions. Sophisticated approaches will be needed to unravel subtle, super-imposed spectral features typically of natural systems, and to maximize the science return of these instruments. Carefully controlled laboratory investigations using homogeneous mineral separates have demonstrated that variations due to solid solution, changes in modal abundances, and the effects of particle size are well understood from a physical basis. In many cases, these variations can be modeled quantitatively using photometric models, mixing approaches, and deconvolution procedures. However, relative to the spectra of individual mineral components, reflectance spectra of rocks and natural surfaces exhibit a reduced spectral contrast. In addition, soils or regolith, which are likely to dominate any natural planetary surface, exhibit spectral properties that have some similarities to the parent materials, but due to weathering and alteration, differences remain that cannot yet be fully recreated in the laboratory or through mixture modeling. A significant challenge is therefore to integrate modeling approaches to derive both lithologic determinations and include the effects of alteration. We are currently conducting laboratory investigations in lithologic modeling to expand upon the basic results of previous analyses with our initial goal to more closely match physical state of natural systems. The effects of alteration are to be considered separately.

Mustard, John F.↗

Attention-Augmented Parametric Kernel Graph Neural Network (APKGNN) for Node Classification

We present a new graph neural network, the Attention-based Parametric-Kernel augmented Graph Neural Network (APKGNN), developed for node classification tasks. Despite extensive work on modeling multi-faceted relationships between connected nodes of a graph, the effect of attention on edge features mapped to relationships has not yet been analyzed through learning representation. This study derives such an attention vector by first calculating node features corresponding to endpoints of an edge and then aggregating these with extracted local intrinsic patches of a given graph to generate augmented local patch vectors. This process uses a parametric kernel based on Gaussian mixture models (GMMs) to embed local neighborhoods of the graph in local patches. The patch vectors then convolve with the above node features to produce an updated node representation. We show that this new learning representation (APKGNN) achieves higher node classification accuracy on tasks - both standard benchmarks (Cora, PubMed, Citeseer) and new experimental short text corpora where nodes correspond to text documents and words. This implementation of the GNN convolution layer outperforms state-of-the-art (SOTA) algorithms, achieving higher training, validation, and test accuracy by a significant margin on three standard benchmark data sets under both SOTA experimental settings and those for new testbeds.

Bose, Avishek↗

Extrapolation of the Rainflow-Counted Load Ranges for Fatigue Assessment of the Wind Turbine's Blades

Wind turbine design standards recommend the use of statistical modeling coupled with extrapolation of the short-term load data to long-term periods for fatigue reliability assessment. However, statistical error and computational expense can limit the accuracy of such approaches. In the case of wind turbine blades, the errors are more significant because of the high material fatigue exponent that makes the damage estimations more sensitive to variations. In addition, due to different excitation sources, the flapwise load range histogram is not unimodal, and thus its statistical modeling is complex. In the present work, we provide three methods for statistical modeling of the flapwise bending moment ranges including a novel approach based on frequency-based separation of the modes. The first two methods are simplified approaches for modeling the most crucial load ranges using unimodal distributions and the third method involves multimodal distribution fitting. The research is based on 3600 10-minute aeroelastic simulations of DTU 10MW case study wind turbine from which a benchmark damage equivalent load (DEL) is calculated. The DEL calculated by each of the three proposed methods is compared to this reference. The results show that the conventional approach based on using 6 seeds as well as using mixture models fitted on the limited data lead to under-conservative results with errors up to 23%. On the other hand, the simplified unimodal approaches provided in this work can provide conservative estimations of the fatigue damage with mean values 5% and 12% higher than the benchmark. However, the variability of the DEL estimates is higher when using unimodal extrapolation of the load ranges, and the data can be conservative by 17.5%. The proposed unimodal fits suggested for modeling and extrapolation of the blade's load ranges provide less errors relatively and most importantly conservative DEL estimations while maintaining computational efficiency.

blade fatigue↗

Speed of Sound Measurements of Binary Mixtures of Difluoromethane (R-32) with 2,3,3,3-Tetrafluoropropene (R-1234yf) or trans-1,3,3,3-Tetrafluoropropene (R-1234ze(E)) Refrigerants

In this article, sound speed data measured using a dual-path pulse-echo instrument are reported for binary mixtures of difluoromethane (R-32) with 2,3,3,3-tetrafluoropropene (R-1234yf) or trans-1,3,3,3-tetrafluoropropene (R-1234ze(E)). The sound speed is reported at two compositions for each binary mixture of approximately (0.33/67) and (0.67/0.33) mole fraction at temperatures between 230 K and 345 K. Data are reported from pressures slightly above the bubble point to 12 MPa for R-32/1234yf mixtures to avoid potential polymerization reactions and to 53 MPa for the R-32/1234ze(E) mixtures. The mean uncertainty of the sound speed data are less than 0.1% of the measured value where uncertainties at individual state points range from 0.04% to 0.5% of the measured value as the conditions approach the mixture critical region. The reported data are compared to available Helmholtz-energy-explicit EOS included in REFPROP and all systems studied have average absolute deviations greater than 2%. The comparisons show that further adjustments to the mixture models are needed to provide a reasonable representation of the data within its experimental uncertainty.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deciphering the Spectra of Flowers to Map Landscape-scale Blooming Dynamics

Like leaves, floral coloration is driven by inherent optical properties, which are determined by pigments, scattering structure, and thickness. However, establishing the relative contribution of these factors to canopy spectral signals is usually limited to in-situ observations. Modeling flowering dynamics (e.g., blooming duration, spatial distribution) at the landscape scale may reveal insights into ecological processes and phenological adaptations to environmental changes. Multitemporal visible to shortwave infrared (VSWIR) imaging spectroscopy observations are especially suited for such efforts. Reflectance in this spectral range is sensitive to major flower pigments, flowering phenology traces, and biophysical differences between flowers and other plant parts. We explored how flowers contribute to spectral signals using a time series of imagery from the Airborne Visible InfraRed Imaging Spectrometer - Next Generation (AVIRIS-NG) collected as part of the SBG High-Frequency Time Series (SHIFT) campaign as a case study. Airborne data were collected weekly during the spring of 2022 across two natural reserves in California. Field spectra were gathered from blooming plots at leaf, flower, and canopy levels at two time points during the campaign. The processed data was used to investigate flowering species' spectro-temporal variation and spatial distribution using Spectral Mixture Residual, Gaussian clustering techniques, and a proposed narrow-band flowering index. Linear spectral unmixing allowed the computation of the weighted contribution of four major high-variance endmembers (leaves, flowers, soil, dark) and low-variance residual signal that comprises subtle spectral features used to track biophysical processes. The reflectance residual was projected on a low principal component basis to characterize flowering clusters' variation and spatial distribution based on the Gaussian mixture model, providing an uncertainty metric to assess the results. Mapping flowering events from modeling spectro-temporal dynamics throughout the season, from pre-blooming to post-flowering stages, allowed us to identify gradient variations in spectral features within the VSWIR spectral range linked to flowering pigments. Time series of the Mixture Residual Blooming Index and the Red-Edge Normalized Difference Vegetation Index revealed specific flowering and greenness phenophases across the two main species (Coreopsis gigantea, Artemisia californica) in the flowering areas. Overall, our approach opens opportunities for future satellite monitoring of floral cycles at broader scales.

Yoseline Angel↗

Flow-based likelihoods for non-Gaussian inference

We investigate the use of data-driven likelihoods to bypass a key assumption made in many scientific analyses, which is that the true likelihood of the data is Gaussian. In particular, we suggest using the optimization targets of flow-based generative models, a class of models that can capture complex distributions by transforming a simple base distribution through layers of nonlinearities. We call these flow-based likelihoods (FBL). We analyze the accuracy and precision of the reconstructed likelihoods on mock Gaussian data, and show that simply gauging the quality of samples drawn from the trained model is not a sufficient indicator that the true likelihood has been learned. We nevertheless demonstrate that the likelihood can be reconstructed to a precision equal to that of sampling error due to a finite sample size. We then apply FBLs to mock weak lensing convergence power spectra, a cosmological observable that is significantly non-Gaussian (NG). We find that the FBL captures the NG signatures in the data extremely well, while other commonly used data-driven likelihoods, such as Gaussian mixture models and independent component analysis, fail to do so. This suggests that works that have found small posterior shifts in NG data with data-driven likelihoods such as these could be underestimating the impact of non-Gaussianity in parameter constraints. By introducing a suite of tests that can capture different levels of NG in the data, we show that the success or failure of traditional data-driven likelihoods can be tied back to the structure of the NG in the data. Here, unlike other methods, the flexibility of the FBL makes it successful at tackling different types of NG simultaneously. Because of this, and consequently their likely applicability across datasets and domains, we encourage their use for inference when sufficient mock data are available for training.

79 ASTRONOMY AND ASTROPHYSICS↗

New Constraints on the Volatile Deposit in Mercury’s North Polar Crater, Prokofiev

We present new high-resolution topographic, illumination, and thermal models of Mercury’s 112 km-diameter north polar crater, Prokofiev. The new models confirm previous results that water ice is stable at the surface within the permanently shadowed region (PSR) of Prokofiev for geologic timescales. The largest radar-bright region in Prokofiev is confirmed to extend up to several kilometers past the boundary of its PSR making it unique on Mercury for hosting a significant radar-bright area outside a PSR. The near-infrared normal albedo distribution of Prokofiev’s PSR suggests the presence of a darkening agent rather than pure surface ice. Linear mixture models predict at least roughly half of the surface area to be covered with this dark material. Using improved altimetry in this crater, we place an upper limit of 26 m on its ice deposit thickness. The 1 km-baseline topographic slope and roughness of the radar-bright deposit are lower than the non-radar-bright floor although the difference is not statistically significant when compared to the non-radar-bright floor’s natural topographic variations. These results place new constraints on the nature of Prokofiev’s volatile deposit that will inform future missions, such as BepiColombo.

Michael K Barker↗

Machine Learning Approaches to Increasing Value of Spaceflight Omics Databases

The number of spaceflight bioscience mission opportunities is too small to allow all relevant biological and environmental parameters to be experimentally identified. Simulated spaceflight experiments in ground-based facilities (GBFs), such as clinostats, are each suitable only for particular investigations -- a rotating-wall vessel may be 'simulated microgravity' for cell differentiation (hours), but not DNA repair (seconds) -- and introduce confounding stimuli, such as motor vibration and fluid shear effects. This uncertainty over which biological mechanisms respond to a given form of simulated space radiation or gravity, as well as its side effects, limits our ability to baseline spaceflight data and validate mission science. Machine learning techniques autonomously identify relevant and interdependent factors in a data set given the set of desired metrics to be evaluated: to automatically identify related studies, compare data from related studies, or determine linkages between types of data in the same study. System-of-systems (SoS) machine learning models have the ability to deal with both sparse and heterogeneous data, such as that provided by the small and diverse number of space biosciences flight missions; however, they require appropriate user-defined metrics for any given data set. Although machine learning in bioinformatics is rapidly expanding, the need to combine spaceflight/GBF mission parameters with omics data is unique. This work characterizes the basic requirements for implementing the SoS approach through the System Map (SM) technique, a composite of a dynamic Bayesian network and Gaussian mixture model, in real-world repositories such as the GeneLab Data System and Life Sciences Data Archive. The three primary steps are metadata management for experimental description using open-source ontologies, defining similarity and consistency metrics, and generating testing and validation data sets. Such approaches to spaceflight and GBF omics data may soon enable unique insight into which measured phenomena correlate to biological mechanisms that are truly affected by spaceflight conditions; which are most likely to be confounded by other variables; and which are insufficiently characterized, significantly increasing existing and future science return from ISS and spaceflight missions.

Gentry, Diana↗