Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multivariate regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Remote quantification of Cm(III) and HNO 3 by fluorescence spectroscopy and chemometrics

A unique approach to remotely quantify Cm(III) (0–100 µg mL −1 ) in HNO 3 (1–12 M) using steady-state laser fluorescence spectroscopy and multivariate regression models was developed. Photoluminescence is amenable to remote measurements using fiber-optic cables and is sensitive to numerous lanthanide and actinide species. In-line measurements can provide feedback to support complex processing in harsh environments (e.g., hot cells) to help guide and optimize radiochemical separations. In this work, Cm(III) spectra were acquired remotely in a glove box as a function of HNO 3 concentration to better understand spectral characteristics and evaluate the utility of multivariate regression models in this system. Furthermore, the Cm(III) fluorescence peak shape, width, position, and intensity changed significantly as a function of HNO 3 concentration, likely because of the displacement of emission quenching inner-sphere water molecules and complexation with nitrate ions. Despite significant covariance and nonlinearity in the data, a D-optimal design strategy successfully minimized training set sample size and was used to build effective partial least squares regression models for Cm(III) and HNO 3 concentrations without a priori knowledge of solution conditions. Chemometrics for modeling complex fluorescence spectra are promising and may find widespread applicability for online analysis in numerous chemical systems found in the nuclear field.

Actinide↗

Predicting major element mineral/melt equilibria - A statistical approach

Empirical equations have been developed for calculating the mole fractions of NaO0.5, MgO, AlO1.5, SiO2, KO0.5, CaO, TiO2, and FeO in a solid phase of initially unknown identity given only the composition of the coexisting silicate melt. The approach involves a linear multivariate regression analysis in which solid composition is expressed as a Taylor series expansion of the liquid compositions. An internally consistent precision of approximately 0.94 is obtained, that is, the nature of the liquidus phase in the input data set can be correctly predicted for approximately 94% of the entries. The composition of the liquidus phase may be calculated to better than 5 mol % absolute. An important feature of this 'generalized solid' model is its reversibility; that is, the dependent and independent variables in the linear multivariate regression may be inverted to permit prediction of the composition of a silicate liquid produced by equilibrium partial melting of a polymineralic source assemblage.

Hostetler, C. J.↗

Probabilistic Modeling of the Renal Stone Formation Module

The Integrated Medical Model (IMM) is a probabilistic tool, used in mission planning decision making and medical systems risk assessments. The IMM project maintains a database of over 80 medical conditions that could occur during a spaceflight, documenting an incidence rate and end case scenarios for each. In some cases, where observational data are insufficient to adequately define the inflight medical risk, the IMM utilizes external probabilistic modules to model and estimate the event likelihoods. One such medical event of interest is an unpassed renal stone. Due to a high salt diet and high concentrations of calcium in the blood (due to bone depletion caused by unloading in the microgravity environment) astronauts are at a considerable elevated risk for developing renal calculi (nephrolithiasis) while in space. Lack of observed incidences of nephrolithiasis has led HRP to initiate the development of the Renal Stone Formation Module (RSFM) to create a probabilistic simulator capable of estimating the likelihood of symptomatic renal stone presentation in astronauts on exploration missions. The model consists of two major parts. The first is the probabilistic component, which utilizes probability distributions to assess the range of urine electrolyte parameters and a multivariate regression to transform estimated crystal density and size distributions to the likelihood of the presentation of nephrolithiasis symptoms. The second is a deterministic physical and chemical model of renal stone growth in the kidney developed by Kassemi et al. The probabilistic component of the renal stone model couples the input probability distributions describing the urine chemistry, astronaut physiology, and system parameters with the physical and chemical outputs and inputs to the deterministic stone growth model. These two parts of the model are necessary to capture the uncertainty in the likelihood estimate. The model will be driven by Monte Carlo simulations, continuously randomly sampling the probability distributions of the electrolyte concentrations and system parameters that are inputs into the deterministic model. The total urine chemistry concentrations are used to determine the urine chemistry activity using the Joint Expert Speciation System (JESS), a biochemistry model. Information used from JESS is then fed into the deterministic growth model. Outputs from JESS and the deterministic model are passed back to the probabilistic model where a multivariate regression is used to assess the likelihood of a stone forming and the likelihood of a stone requiring clinical intervention. The parameters used to determine to quantify these risks include: relative supersaturation (RS) of calcium oxalate, citrate/calcium ratio, crystal number density, total urine volume, pH, magnesium excretion, maximum stone width, and ureteral location. Methods and Validation: The RSFM is designed to perform a Monte Carlo simulation to generate probability distributions of clinically significant renal stones, as well as provide an associated uncertainty in the estimate. Initially, early versions will be used to test integration of the components and assess component validation and verification (V&V), with later versions used to address questions regarding design reference mission scenarios. Once integrated with the deterministic component, the credibility assessment of the integrated model will follow NASA STD 7009 requirements.

kidney stones↗

Modeling of topographic effects on Antarctic sea ice using multivariate adaptive regression splines

The role of seafloor topography in the spatial variations of the southern ocean sea ice cover as observed (every other day) by the Nimbus 7 scanning multichannel microwave radiometer satellite in the years 1980, 1983, and 1984 is studied. Bottom bathymetry can affect sea ice surface characteristics because of the basically barotropic circulation of the ocean south of the Antarctic Circumpolar current. The main statistical tool used to quantify this effect is a local nonparametric regression model of sea ice concentration as a function of the depth and its first two derivatives in both meridional and zonal directions. First, we model the relationship of bathymetry to sea ice concentration in two sudy areas, one over the Maud Rise and the other over the Ross Sea shelf region. The multiple correlation coefficient is found to average 44% in the Maud Rise study area and 62% in the Ross Sea study area over the years 1980, 1983, and 1984. Second, a strategy of dividing the entire Antarctic region into an overlapping mosaic of small areas, or windows is considered. Keeping the windows small reduces the correlation of bathymetry with other factors such as wind, sea temperature, and distance to the continent. We find that although the form of the model varies from window to window due to the changing role of other relevant environmental variables, we are left with a spatially consistent ordering of the relative importance of the topographic predictors. For a set of three representative days in the Austral winter of 1980, the analysis shows that an average of 54% of the spatial variation in sea ice concentration over the entire ice cover can be attributed to topographic variables. The results thus support the hypothesis that there is a sea ice to bottom bathymetry link. However this should not undermine the considerable influence of wind, current, and temperature which affect the ice distribution directly and are partly responsible for the observed bathymetric effects.

De Veaux, Richard D.↗

An Alternative Flight Software Trigger Paradigm: Applying Multivariate Logistic Regression to Sense Trigger Conditions using Inaccurate or Scarce Information

In late 2014, NASA will fly the Orion capsule on a Delta IV-Heavy rocket for the Exploration Flight Test-1 (EFT-1) mission. For EFT-1, the Orion capsule will be flying with a new GPS receiver and new navigation software. Given the experimental nature of the flight, the flight software must be robust to the loss of GPS measurements. Once the high-speed entry is complete, the drogue parachutes must be deployed within the proper conditions to stabilize the vehicle prior to deploying the main parachutes. When GPS is available in nominal operations, the vehicle will deploy the drogue parachutes based on an altitude trigger. However, when GPS is unavailable, the navigated altitude errors become excessively large, driving the need for a backup barometric altimeter. In order to increase overall robustness, the vehicle also has an alternate method of triggering the drogue parachute deployment based on planet-relative velocity if both the GPS and the barometric altimeter fail. However, this velocity-based trigger results in large altitude errors relative to the targeted altitude. Motivated by this challenge, this paper demonstrates how logistic regression may be employed to automatically generate robust triggers based on statistical analysis. Logistic regression is used as a ground processor pre-flight to develop a classifier. The classifier would then be implemented in flight software and executed in real-time. This technique offers excellent performance even in the face of highly inaccurate measurements. Although the logistic regression-based trigger approach will not be implemented within EFT-1 flight software, the methodology can be carried forward for future missions and vehicles.

Smith, Kelly M.↗

An Alternative Flight Software Trigger Paradigm: Applying Multivariate Logistic Regression to Sense Trigger Conditions Using Inaccurate or Scarce Information

In late 2014, NASA will fly the Orion capsule on a Delta IV-Heavy rocket for the Exploration Flight Test-1 (EFT-1) mission. For EFT-1, the Orion capsule will be flying with a new GPS receiver and new navigation software. Given the experimental nature of the flight, the flight software must be robust to the loss of GPS measurements. Once the high-speed entry is complete, the drogue parachutes must be deployed within the proper conditions to stabilize the vehicle prior to deploying the main parachutes. When GPS is available in nominal operations, the vehicle will deploy the drogue parachutes based on an altitude trigger. However, when GPS is unavailable, the navigated altitude errors become excessively large, driving the need for a backup barometric altimeter to improve altitude knowledge. In order to increase overall robustness, the vehicle also has an alternate method of triggering the parachute deployment sequence based on planet-relative velocity if both the GPS and the barometric altimeter fail. However, this backup trigger results in large altitude errors relative to the targeted altitude. Motivated by this challenge, this paper demonstrates how logistic regression may be employed to semi-automatically generate robust triggers based on statistical analysis. Logistic regression is used as a ground processor pre-flight to develop a statistical classifier. The classifier would then be implemented in flight software and executed in real-time. This technique offers improved performance even in the face of highly inaccurate measurements. Although the logistic regression-based trigger approach will not be implemented within EFT-1 flight software, the methodology can be carried forward for future missions and vehicles.

Smith, Kelly M.↗

An Alternative Flight Software Paradigm: Applying Multivariate Logistic Regression to Sense Trigger Conditions using Inaccurate or Scarce Information

In late 2014, NASA will fly the Orion capsule on a Delta IV-Heavy rocket for the Exploration Flight Test-1 (EFT-1) mission. For EFT-1, the Orion capsule will be flying with a new GPS receiver and new navigation software. Given the experimental nature of the flight, the flight software must be robust to the loss of GPS measurements. Once the high-speed entry is complete, the drogue parachutes must be deployed within the proper conditions to stabilize the vehicle prior to deploying the main parachutes. When GPS is available in nominal operations, the vehicle will deploy the drogue parachutes based on an altitude trigger. However, when GPS is unavailable, the navigated altitude errors become excessively large, driving the need for a backup barometric altimeter to improve altitude knowledge. In order to increase overall robustness, the vehicle also has an alternate method of triggering the parachute deployment sequence based on planet-relative velocity if both the GPS and the barometric altimeter fail. However, this backup trigger results in large altitude errors relative to the targeted altitude. Motivated by this challenge, this paper demonstrates how logistic regression may be employed to semi-automatically generate robust triggers based on statistical analysis. Logistic regression is used as a ground processor pre-flight to develop a statistical classifier. The classifier would then be implemented in flight software and executed in real-time. This technique offers improved performance even in the face of highly inaccurate measurements. Although the logistic regression-based trigger approach will not be implemented within EFT-1 flight software, the methodology can be carried forward for future missions and vehicles

Smith, Kelly↗

Causal correlation of foliar biochemical concentrations with AVIRIS spectra using forced entry linear regression

A major goal of airborne imaging spectrometry is to estimate the biochemical composition of vegetation canopies from reflectance spectra. Remotely-sensed estimates of foliar biochemical concentrations of forests would provide valuable indicators of ecosystem function at regional and eventually global scales. Empirical research has shown a relationship exists between the amount of radiation reflected from absorption features and the concentration of given biochemicals in leaves and canopies (Matson et al., 1994, Johnson et al., 1994). A technique commonly used to determine which wavelengths have the strongest correlation with the biochemical of interest is unguided (stepwise) multiple regression. Wavelengths are entered into a multivariate regression equation, in their order of importance, each contributing to the reduction of the variance in the measured biochemical concentration. A significant problem with the use of stepwise regression for determining the correlation between biochemical concentration and spectra is that of 'overfitting' as there are significantly more wavebands than biochemical measurements. This could result in the selection of wavebands which may be more accurately attributable to noise or canopy effects. In addition, there is a real problem of collinearity in that the individual biochemical concentrations may covary. A strong correlation between the reflectance at a given wavelength and the concentration of a biochemical of interest, therefore, may be due to the effect of another biochemical which is closely related. Furthermore, it is not always possible to account for potentially suitable waveband omissions in the stepwise selection procedure. This concern about the suitability of stepwise regression has been identified and acknowledged in a number of recent studies (Wessman et al., 1988, Curran, 1989, Curran et al., 1992, Peterson and Hubbard, 1992, Martine and Aber, 1994, Kupiec, 1994). These studies have pointed to the lack of a physical link between wavelengths chosen by stepwise regression and the biochemical of interest, and this in turn has cast doubts on the use of imaging spectrometry for the estimation of foliar biochemical concentrations at sites distant from the training sites. To investigate this problem, an analysis was conducted on the variation in canopy biochemical concentrations and reflectance spectra using forced entry linear regression.

Dawson, Terence P.↗

Responsiveness of miscanthus and switchgrass yields to stand age and nitrogen fertilization: A meta‐regression analysis

Abstract Optimal management of the perennial bioenergy crops, miscanthus and switchgrass, requires an understanding of their responsiveness to nitrogen (N) fertilizer at different maturity stages across locations and growing conditions. Earlier studies that have examined the yield response of these crops to N and stand age using field experiments or meta‐analysis techniques provide mixed evidence. We extend earlier studies by applying a multi‐level mixed‐effects (MLME) meta‐regression model to conduct a more extensive multivariate regression of yield response of these crops to N and stand age, while controlling for climate and location conditions and unobserved factors related to study design. Our findings are based on 1403 and 2811 yield observations for miscanthus and switchgrass, respectively, from experiments conducted between 2002 and 2019 across the rainfed region in the United States. We find statistically significant evidence that an additional year of maturity increases miscanthus and switchgrass yields but at a decreasing rate; yields peak at the 7th and 6th year respectively, for the observed range of applied N rates and stands. We also find that an increase in N application increases yield by a statistically significant level, but at a declining rate; the magnitude of the yield response to N is, however, small and varies with the age of the crop. The impact of N is larger on older compared to younger and middle‐aged stands of miscanthus. In contrast, the impact of N on switchgrass is larger on middle‐aged compared to younger and older stands of switchgrass. We do not find a statistically significant effect of soil productivity on yield for either crop. This analysis provides a basis for developing N application recommendations and optimal rotation age for miscanthus and switchgrass and shows that these energy crops can grow just as productively on low productivity land as on high productivity land.

59 BASIC BIOLOGICAL SCIENCES↗

Post-landing major element quantification using SuperCam laser induced breakdown spectroscopy

The SuperCam instrument on the Perseverance Mars 2020 rover uses a pulsed 1064 nm laser to ablate targets at a distance and conduct laser induced breakdown spectroscopy (LIBS) by analyzing the light from the resulting plasma. SuperCam LIBS spectra are preprocessed to remove ambient light, noise, and the continuum signal present in LIBS observations. Prior to quantification, spectra are masked to remove noisier spectrometer regions and spectra are normalized to minimize signal fluctuations and effects of target distance. In some cases, the spectra are also standardized or binned prior to quantification. To determine quantitative elemental compositions of diverse geologic materials at Jezero crater, Mars, we use a suite of 1198 laboratory spectra of 334 well-characterized reference samples. The samples were selected to span a wide range of compositions and include typical silicate rocks, pure minerals (e.g., silicates, sulfates, carbonates, oxides), more unusual compositions (e.g., Mn ore and sodalite), and replicates of the sintered SuperCam calibration targets (SCCTs) onboard the rover. For each major element (SiO 2 , TiO 2 , Al 2 O 3 , FeO T , MgO, CaO, Na 2 O, K 2 O), the database was subdivided into five “folds” with similar distributions of the element of interest. One fold was held out as an independent test set, and the remaining four folds were used to optimize multivariate regression models relating the spectrum to the composition. We considered a variety of models, and selected several for further investigation for each element, based primarily on the root mean squared error of prediction (RMSEP) on the test set, when analyzed at 3 m. In cases with several models of comparable performance at 3 m, we incorporated the SCCT performance at different distances to choose the preferred model. Shortly after landing on Mars and collecting initial spectra of geologic targets, we selected one model per element. Subsequently, with additional data from geologic targets, some models were revised to ensure results that are more consistent with geochemical constraints. The calibration discussed here is a snapshot of an ongoing effort to deliver the most accurate chemical compositions with SuperCam LIBS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Risk Ratio and Risk Difference Estimation in Case-cohort Studies

Background: In case-cohort studies with binary outcomes, ordinary logistic regression analyses have been widely used because of their computational simplicity. However, the resultant odds ratio estimates cannot be interpreted as relative risk measures unless the event rate is low. The risk ratio and risk difference are more favorable outcome measures that are directly interpreted as effect measures without the rare disease assumption. Methods: We provide pseudo-Poisson and pseudo-normal linear regression methods for estimating risk ratios and risk differences in analyses of case-cohort studies. These multivariate regression models are fitted by weighting the inverses of sampling probabilities. Also, the precisions of the risk ratio and risk difference estimators can be improved using auxiliary variable information, specifically by adapting the calibrated or estimated weights, which are readily measured on all samples from the whole cohort. Finally, we provide computational code in R (R Foundation for Statistical Computing, Vienna, Austria) that can easily perform these methods. Results: Through numerical analyses of artificially simulated data and the National Wilms Tumor Study data, accurate risk ratio and risk difference estimates were obtained using the pseudo-Poisson and pseudo-normal linear regression methods. Also, using the auxiliary variable information from the whole cohort, precisions of these estimators were markedly improved. Conclusion: The ordinary logistic regression analyses may provide uninterpretable effect measure estimates, and the risk ratio and risk difference estimation methods are effective alternative approaches for case-cohort studies. These methods are especially recommended under situations in which the event rate is not low.

60 APPLIED LIFE SCIENCES↗

Regression Model Optimization for the Analysis of Experimental Data

A candidate math model search algorithm was developed at Ames Research Center that determines a recommended math model for the multivariate regression analysis of experimental data. The search algorithm is applicable to classical regression analysis problems as well as wind tunnel strain gage balance calibration analysis applications. The algorithm compares the predictive capability of different regression models using the standard deviation of the PRESS residuals of the responses as a search metric. This search metric is minimized during the search. Singular value decomposition is used during the search to reject math models that lead to a singular solution of the regression analysis problem. Two threshold dependent constraints are also applied. The first constraint rejects math models with insignificant terms. The second constraint rejects math models with near-linear dependencies between terms. The math term hierarchy rule may also be applied as an optional constraint during or after the candidate math model search. The final term selection of the recommended math model depends on the regressor and response values of the data set, the user s function class combination choice, the user s constraint selections, and the result of the search metric minimization. A frequently used regression analysis example from the literature is used to illustrate the application of the search algorithm to experimental data.

Ulbrich, N.↗

Spatially Refined Satellite Gravimetry Captures Human Signatures in Global Terrestrial Water Storage Trends

Human activities have directly altered the water cycle through water management, aquifer pumping, agricultural irrigation, and land use change. Although satellite gravimetry has transformed global hydrological research, its coarse resolution limits attribution of freshwater change to human activities at many management-relevant scales. Here we assessed global terrestrial water storage (TWS) trends from April 2002 to November 2025 using “stacked” regression of Level-1B intersatellite ranging data, which leverages temporal information and variability to dramatically improve effective spatial resolution relative to standard approaches. We combined this refined product with rigorous uncertainty analysis, autocorrelation-robust geostatistical methods, and literature assessment to evaluate TWS trend associations with land and water use, climate variability, and glacial mass loss. We identified TWS trend hotspots exhibiting significant spatial associations with anthropogenic 40 drivers, including groundwater and surface-water irrigation, rainfed agriculture, deforestation, and reservoir impoundment. Across these regions, cumulative TWS losses (3,122 Gt) substantially exceeded gains (2,432 Gt). Compared with traditional regression of monthly mascons, our approach yielded regional trend magnitudes that are on average 33% larger, revealing that global freshwater depletion, particularly from groundwater pumping, is considerably more acute than previously estimated. Multivariate regression models show that humans account for a significant share of the spatial variability in TWS trends on every non-polar continent except Australia. We detected localized TWS gains linked to rainfed agriculture, surface water irrigation, and reservoir filling that were unresolved in earlier gravimetric studies. The methodology provides a foundation for future gravity missions to independently track decadal freshwater change with unprecedented spatial fidelity.

groundwater↗

Monitoring Sulfuric Acid and Temperature Using Raman Spectroscopy and Multivariate Chemometrics

Multivariate regression models were optimized for the quantification of sulfuric acid (H 2 SO 4 ) [0–8 M] and temperature (20 °C–80 °C) in the presence of ammonium sulfate ((NH 4 ) 2 SO 4 [0–0.6 M]) using Raman spectroscopy. Optical vibrational spectroscopy is a useful nondestructive technique for the in situ analysis of complex chemical systems notoriously difficult to monitor in situ and in real-time. Multivariate analysis, a chemometrics method, can be paired with these nondestructive optical methods for determining analyte concentration and speciation in complex solutions, such as dissociated species in polyprotic acids, e.g., H 2 SO 4 . The effect of temperature is often overlooked although it can have a major influence on speciation and the corresponding Raman spectra. Here, in this study, partial least squares regression models were optimized for the quantification of H 2 SO 4 and its two deprotonated forms as a function of temperature. Measuring bisulfate as a function of temperature is particularly challenging owing to changes in the second dissociation constant. A designed training set effectively minimized the sample set size and trained a robust predictive model with percent root mean square error of <3% for H 2 SO 4 . The practical strategy employed here was demonstrated to be effective for building chemometric models that directly account for dynamic temperatures with static samples and is shown to be amenable to flow cell analysis applications with a simple calibration transfer for process monitoring applications.

D-optimal design↗

Simultaneous quantification of uranium( VI ), samarium, nitric acid, and temperature with combined ensemble learning, laser fluorescence, and Raman scattering for real-time monitoring

In this work, laser-induced fluorescence spectroscopy (LIFS), Raman spectroscopy, and a stacked regression ensemble was developed for near real-time quantification of uranium(VI) (1–100 μg mL –1 ), samarium (0–200 μg mL –1 ) and nitric acid (0.1–4 M) with varying temperature (20 °C–45 °C). LIFS applications range from fundamental lab-scale studies to real-time process monitoring at industrial levels, such as nuclear reprocessing applications, provided the phenomena affecting the fluorescence spectrum are accounted for (e.g., absorption, quenching, complexation). Multiple chemometric models were examined and compared to a more traditional multivariate regression approach called partial least squares (PLS). Results obtained on synthetic samples selected using D-optimal experimental design indicated that a stacked regression method, which included ridge regression, random forest, PLS, and an eXtreme gradient boost algorithm, successfully measured uranium(VI) concentrations directly in nitric acid without measuring luminescence lifetimes or standard addition. The top model resulted in percent root-mean-square error of prediction values of 5.2, 1.9, 3.0, and 2.3% for U(VI), Sm 3+ , HNO 3 , and temperature, respectively. The approach may be useful for quantifying fluorescent fission products (e.g., Sm 3+ ) to provide information on burnup of irradiated nuclear fuel. This novel framework reinforces the applicability of LIFS for real-time applications in nuclear fuel cycle applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Probabilistic Modeling of a Three-Stage Human Landing System Architecture

Space Policy Directive-1 has led to NASA partnerships with commercial entities on procurement which includes the development of the Human Landing System (HLS) [1]. With the goal of delivering human crew to the lunar surface by 2024, system uncertainties become an important obstacle to the maturation of multiple new, driving technologies and mission concepts of the HLS program. As unmitigated uncertainties have previously led to failed development programs, these risks and their impacts must be understood and handled to ensure program success [2]. Sources of uncertainty include novel engine designs and configurations, increased reliance on cryogenic fluid management(CFM), and refueling technologies—which propagate as high-level performance metrics such as overall propellant mass and engine performance. Also, the occurrence of operational uncertainties—e.g. launch conditions or need to abort during the mission—can cause cascading effects on the rest of the mission that are difficult to definitively quantify, and are outside the scope of control. These concrete examples and other occurrences can be categorized as either epistemic or aleatory uncertainties.Epistemic uncertainty arises due to a lack of knowledge and can be alleviated with design and program maturation. Aleatory uncertainty is due to the inherent randomness of the system and cannot be directly reduced, unlike epistemic uncertainty. Robust design and probabilistic methods can compensate for aleatory effects. A taxonomy of uncertainty is referred to for this work [3]. In this paper, a probabilistic methodology to handle uncertainties has been demonstrated on a three-element HLS concept [1, 4], which allows tracking of current best estimates of the concept and assessment of concept design robustness against uncertainties. A sample case has been completed for this abstract, and an expansion on the methodology will be included in the final paper. This methodology has two key parts: first, the creation of a dynamic architecture model of a three-element HLS concept; and second, its use with surrogate modeling and range estimating techniques to capture and propagate uncertainties. This abstract will cover the basics of the approach used, and further details and justifications will be in the final paper.The mission profile associated with this three-element concept (Fig 1) was modeled as a set of mission events that facilitated mass changes, idles, or spacecraft maneuvers. The mission profile scope starts with each element’s NRHO orbit insertion and aggregation and ends at post-sortie rendezvous with Orion. More detail on the mission profile will be in the final paper. The DYnamic Rocket EQuation Tool (DYREQT), a space systems synthesis and sizing framework used by NASA, was used as the physics framework to model the HLS architecture for applying the probabilistic methodology [5, 6]. Specifically, a parametric representation of the lander, ascent, and transfer elements and the mission profile of each element was established, with vehicle and mission parameters available as inputs to allow for a dynamic model. Each vehicle stage was modeled with high-level performance metrics, using Isp and propellant mass fraction (PMF) to remain parametric. For the probabilistic analysis, uncertainties of interest within the HLS concept were enumerated and represented as parameters within the DYREQT model as inputs for vehicle stages or mission profile events. These parameters were frozen at their nominal values for the purposes of baselining architecture performance and sizing the vehicle appropriately based on reference documentation [1]. Range estimating—a probabilistic method that combines Monte Carlo sampling, focus on critical parameters, and heuristics to assess risk and opportunities—is traditionally used with Mass Equipment Lists (MELs), but has been adapted with operational parameters as well as vehicle parameters in theDYREQT model to capture mission uncertainty alongside vehicle uncertainty [7, 3]. This method was selected due to its application and insight on a system from a bottom-up perspective, independence from historical rules of thumb, and ability to generate sensitivities based on design decisions and uncertainties. As a sample case for the abstract, the boiloff rates of the vehicle elements and the loiter times during the mission (simulating launch time variations and changing window of opportunities) were used with range estimating to provide preliminary results. To perform the range estimation portion of this methodology (depicted in Fig. 3, further details in final paper), the DYREQT model was sampled using a Design of Experiments (DoE) to efficiently explore the architecture design space with respect to the sample set of uncertainty parameters; 5,000 cases via Latin Hypercube Sampling were computed on the DYREQT architecture model. Then, the results were used to create surrogate models, multivariate regressions that can visualize hypercube trends in the design space, of the architecture with respect to the uncertainty parameters. Range estimating was applied to the surrogates instead of the actual models, which saves computational expense due to the bulk of cases needed for the Monte Carlo simulation as part of range estimating. Uncertainty parameters were sampled independently from triangular distributions using the DoE ranges as ‘min’ and ‘max’, and the nominal value as ‘most likely’. Based engineering intuition, some uncertainty parameters are correlated—e.g. if the main propellant has a high boil-off rate, the oxidizer should follow suit as both are related to CFM technology.While a Monte Carlo simulation samples all inputs as independent, the results would show model correlations; thus, it is efficient to sample the inputs as correlated. Using a correlation matrix constructed for the uncertainty parameters, previously independent samples were transformed to perform a Correlated Monte Carlo. A table for the DoE ranges and probability distribution parameters is shown in Table 1, and more details on Correlated Monte Carlo Simulations will be discussed in the final paper. The model’s resulting DoE showed that multivariate polynomial equations fit via least squares method captured its behavior accurately for the sample case. For the Correlated Monte Carlo Simulation, a positive correlation between fuel and oxidizer boiloff rates was used as a demonstration. 10,000 cases were computed with the surrogates and the launched masses for each vehicle element was collated. The results can be displayed in a probability density function (PDF), showing the impact of the uncertainty parameters chosen. Integrating the PDFs will yield a cumulative distribution function (CDF) that shows the cumulative probability of a given value on the x-axis. For the sample case, the elements’ launch mass margin was calculated and represented in as CDFs, as a demonstrated representation of figures of merit for the HLS concept. For the lander and ascent elements, the NRHO mass insertion limit is 16t; the transfer element has a limit of 30t [1]. It can be seen with Figure 2 that this probabilistic methodology can provide insight into mass margin with respect to the uncertainties being modeled. Currently, the results show that the lander (descent) vehicle element has the most restrictive design space; it is the only element to show a 10% probability of negative margin. Further analysis on the Monte Carlo results will show sensitivities for driving constraints and parameters for architecture feasibility, which can lead to establishing potential mission rules.The combination of range estimating with a parametric architecture model for HLS demonstrated the capability of this probabilistic methodology in a sample case. As the HLS development progresses, this methodology has the potential for keeping current best estimates of architecture performance for awarded concepts due to the flexibility in DYREQT’s modeling framework and its parametric nature. Concept maturation and increased epistemic knowledge can be injected into the model probabilistic modeling, and thus continue to track probability of mission success.

Stephanie Y Zhu↗

NCAPH drives breast cancer progression and identifies a gene signature that predicts luminal a tumour recurrence

Luminal A tumours generally have a favourable prognosis but possess the highest 10-year recurrence risk among breast cancers. Additionally, a quarter of the recurrence cases occur within 5 years post-diagnosis. Identifying such patients is crucial as long-term relapsers could benefit from extended hormone therapy, while early relapsers might require more aggressive treatment. We conducted a study to explore non-structural chromosome maintenance condensin I complex subunit H’s (NCAPH) role in luminal A breast cancer pathogenesis, both in vitro and in vivo, aiming to identify an intratumoural gene expression signature, with a focus on elevated NCAPH levels, as a potential marker for unfavourable progression. Our analysis included transgenic mouse models overexpressing NCAPH and a genetically diverse mouse cohort generated by backcrossing. A least absolute shrinkage and selection operator (LASSO) multivariate regression analysis was performed on transcripts associated with elevated intratumoural NCAPH levels. We found that NCAPH contributes to adverse luminal A breast cancer progression. The intratumoural gene expression signature associated with elevated NCAPH levels emerged as a potential risk identifier. Transgenic mice overexpressing NCAPH developed breast tumours with extended latency, and in Mouse Mammary Tumor Virus (MMTV)-NCAPH ErbB2 double-transgenic mice, luminal tumours showed increased aggressiveness. High intratumoural Ncaph levels correlated with worse breast cancer outcome and subpar chemotherapy response. A 10-gene risk score, termed Gene Signature for Luminal A 10 (GSLA10), was derived from the LASSO analysis, correlating with adverse luminal A breast cancer progression. The GSLA10 signature outperformed the Oncotype DX signature in discerning tumours with unfavourable outcomes, previously categorised as luminal A by Prediction Analysis of Microarray 50 (PAM50) across three independent human cohorts. This new signature holds promise for identifying luminal A tumour patients with adverse prognosis, aiding in the development of personalised treatment strategies to significantly improve patient outcomes.

60 APPLIED LIFE SCIENCES↗

Field evaluation of semi‐automated moisture estimation from geophysics using machine learning

Geophysical methods can provide three-dimensional (3D), spatially continuous estimates of soil moisture. However, point-to-point comparisons of geophysical properties to measure soil moisture data are frequently unsatisfactory, resulting in geophysics being used for qualitative purposes only. This is because (1) geophysics requires models that relate geophysical signals to soil moisture, (2) geophysical methods have potential uncertainties resulting from smoothing and artifacts introduced from processing and inversion, and (3) results from multiple geophysical methods are not easily combined within a single soil moisture estimation framework. To investigate these potential limitations, an irrigation experiment was performed wherein soil moisture was monitored through time, and several surface geophysical datasets indirectly sensitive to soil moisture were collected before and after irrigation: ground penetrating radar, electrical resistivity tomography (ERT), and frequency domain electromagnetics (FDEM). Data were exported in both raw and processed form, and then snapped to a common 3D grid to facilitate moisture prediction by standard calibration techniques, multivariate regression, and machine learning. A combination of inverted ERT data, raw FDEM, and inverted FDEM data was most informative for predicting soil moisture using a random regression forest model (one-thousand 60/40 training/test cross-validation folds produced root mean squared errors ranging from 0.025–0.046 cm 3 /cm 3 ). This cross-validated model was further supported by a separate evaluation using a test set from a physically separate portion of the study area. Machine learning was conducive to a semi-automated model-selection process that could be used for other sites and datasets to locally improve accuracy.

54 ENVIRONMENTAL SCIENCES↗