Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Unique signatures of synoptic features in Tiros N satellite data

The application of satellite data to the study of synoptic characteristics is analyzed. A radiative transfer model is used to generate a database of satellite information on synoptic features. Canonical discriminant analysis is employed to reveal the differences among synoptic sounding classes; and the fine structures within each sounding class are examined with rotated factor analysis. Diagrams of wind and frontal inversions are presented. It is noted that the applicability of satellite data depends on the method used to analyze it and multivariable statistical techniques may be useful for deriving additional information from satellite data.

White, G. A., III↗

Landmark-Warped Emulators for Models with Misaligned Functional Response

Many computer models output functional data, and in some cases, these functional data have similar, but misaligned, shape characteristics. In this paper, we introduce a general approach for building emulators for computer models that output misaligned functional data when key values in the functional response (landmarks) can be easily identified. This approach has two main parts: modeling the aligned (using the landmarks) functional data, and modeling the functions that map the misaligned data to the aligned space (warping functions). As the warping functions are required to be monotonic, we give special attention to modeling monotonic functional response data. We discuss how our approach can be easily applied for a variety of typical emulators, such as Gaussian processes, Bayesian multivariate adaptive regression splines, and Bayesian additive regression trees, and how sensitivity analysis can be performed. We demonstrate our approach by building emulators for two applications: (1) a high-energy-density physics computer model used to simulate inertial confinement fusion ignition experiments, where model outputs are highly misaligned, and (2) a multiphysics continuum hydrocode used to simulate high-velocity impact experiments, where model outputs are only slightly misaligned. In case (1) traditional methods cannot be applied, while in (2) they can be applied, but the proposed method performs significantly better.

97 MATHEMATICS AND COMPUTING↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Comparison of Multivariate Time Series Prediction Techniques for Emulating Noah-LSM Soil Moisture Outputs

Land surface models are crucial tools for many earth science applications including numerical weather prediction, water resource and crop monitoring, and climatological analysis. Given a set of atmospheric forcings, seasonal data, and static parameters, models like Noah-LSM solve for land surface quantities including skin temperature, sensible heat flux, and soil moisture. While these calculations are theoretically robust, they are often computationally expensive. Since artificial neural networks (ANNs) are universal function approximators, they can learn to emulate the output of a deterministic numerical model given a time series of input forcings, with the learned ANN having substantially shorter execution time. The ANN could efficiently parameterize other models, generate ensembles, and provide first-guess inputs for retrievals. As such, with the goal of developing a model that efficiently mimics the output of Noah-LSM given NLDAS2 forcings on a region covering much of the central US, we examine and compare several neural network architectures for the multi-horizon multivariate time series forecasting problem. Recent literature includes a diverse set of approaches including autoregressive architectures like LSTM and GRU, parametric and non-parametric statistical predictors (ForecastNet and MQRNN), self-attention (LSTM-attention-LSTM), and temporal convovlution (DeepTCN). We implement several of these models for the Noah-LSM prediction task, highlighting the features and challenges for each and providing practical insight on the training process.

Mitchell Dodson↗

Physics-Infused AI/ML Based Digital-Twin Framework for Flow-Induced-Vibration Damage Prediction in a Nuclear Reactor Heat Exchanger

This report summarizes some of the ongoing work related to the development of an expert-elicitation-digital-twin framework for real time damage state prediction in heat exchanger components of a nuclear reactor. The framework is targeted towards predicting damage associated with coupled low cycle fatigue (associated with regular heat-up, cool-down and power operation transients) and high cycle fatigue (associated with flow induced vibration transients). The overall framework will be based on a NoSQL based database, physics-infused-geometry-dependent virtual-sensor data, different AI/ML techniques-based data-driven-predictive-model applications (Apps) and real-time plant sensor measurements available through few existing sensors. Towards this overall goal, this report updates some of the ongoing work, such as on implementation of a NoSQL Database (such as MongoDB), FE based heat transfer analysis of a heat exchanger (e.g. of a PWR steam generator) for generating geometry-dependent virtual sensor data and evaluation of various AI/ML models such as based on multivariate linear regression, ensembled decision-tree based Random-Forest and Gradient-Boosting regression and high-dimensional-kernel-function-transformation based Support-Vector-Machine regression models. The AI/ML models were evaluated for predicting multi-time-series thermal states at thousands of 3D point-clouds

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

LinkWinds, A System for Interactive Scientific Data Analysis and Visualization

The Linked Windows Interactive Data System is a prototype visual data exploration system resulting from a NASA/JPL program of research into graphical methods for rapidly and interactively accessing, displaying and analyzing large multivariate multidisciplinary data sets.

dynamic interconnection data-linking paradigm easy↗

The Statistical Consulting Center for Astronomy (SCCA)

The process by which raw astronomical data acquisition is transformed into scientifically meaningful results and interpretation typically involves many statistical steps. Traditional astronomy limits itself to a narrow range of old and familiar statistical methods: means and standard deviations; least-squares methods like chi(sup 2) minimization; and simple nonparametric procedures such as the Kolmogorov-Smirnov tests. These tools are often inadequate for the complex problems and datasets under investigations, and recent years have witnessed an increased usage of maximum-likelihood, survival analysis, multivariate analysis, wavelet and advanced time-series methods. The Statistical Consulting Center for Astronomy (SCCA) assisted astronomers with the use of sophisticated tools, and to match these tools with specific problems. The SCCA operated with two professors of statistics and a professor of astronomy working together. Questions were received by e-mail, and were discussed in detail with the questioner. Summaries of those questions and answers leading to new approaches were posted on the Web (www.state.psu.edu/ mga/SCCA). In addition to serving individual astronomers, the SCCA established a Web site for general use that provides hypertext links to selected on-line public-domain statistical software and services. The StatCodes site (www.astro.psu.edu/statcodes) provides over 200 links in the areas of: Bayesian statistics; censored and truncated data; correlation and regression, density estimation and smoothing, general statistics packages and information; image analysis; interactive Web tools; multivariate analysis; multivariate clustering and classification; nonparametric analysis; software written by astronomers; spatial statistics; statistical distributions; time series analysis; and visualization tools. StatCodes has received a remarkable high and constant hit rate of 250 hits/week (over 10,000/year) since its inception in mid-1997. It is of interest to scientists both within and outside of astronomy. The most popular sections are multivariate techniques, image analysis, and time series analysis. Hundreds of copies of the ASURV, SLOPES and CENS-TAU codes developed by SCCA scientists were also downloaded from the StatCodes site. In addition to formal SCCA duties, SCCA scientists continued a variety of related activities in astrostatistics, including refereeing of statistically oriented papers submitted to the Astrophysical Journal, talks in meetings including Feigelson's talk to science journalists entitled "The reemergence of astrostatistics" at the American Association for the Advancement of Science meeting, and published papers of astrostatistical content.

Akritas, Michael↗

Preparing for the next pandemic via transfer learning from existing diseases with hierarchical multi-modal BERT: a study on COVID-19 outcome prediction

Abstract Developing prediction models for emerging infectious diseases from relatively small numbers of cases is a critical need for improving pandemic preparedness. Using COVID-19 as an exemplar, we propose a transfer learning methodology for developing predictive models from multi-modal electronic healthcare records by leveraging information from more prevalent diseases with shared clinical characteristics. Our novel hierarchical, multi-modal model ( $${\textsc {TransMED}}$$ T R A N S MED ) integrates baseline risk factors from the natural language processing of clinical notes at admission, time-series measurements of biomarkers obtained from laboratory tests, and discrete diagnostic, procedure and drug codes. We demonstrate the alignment of $${\textsc {TransMED}}$$ T R A N S MED ’s predictions with well-established clinical knowledge about COVID-19 through univariate and multivariate risk factor driven sub-cohort analysis. $${\textsc {TransMED}}$$ T R A N S MED ’s superior performance over state-of-the-art methods shows that leveraging patient data across modalities and transferring prior knowledge from similar disorders is critical for accurate prediction of patient outcomes, and this approach may serve as an important tool in the early response to future pandemics.

59 BASIC BIOLOGICAL SCIENCES↗

Multivariate optimum interpolation of surface pressure and surface wind over oceans

The present multivariate analysis method for surface pressure and winds incorporates ship wind observations into the analysis of surface pressure. For the specific case of 0000 GMT, on February 3, 1979, the additional data resulted in a global rms difference of 0.6 mb; individual maxima as larse as 5 mb occurred over the North Atlantic and East Pacific Oceans. These differences are noted to be smaller than the analysis increments to the first-guess fields.

Bloom, S. C.↗

Inferring Instantaneous, Multivariate and Nonlinear Sensitivities for the Analysis of Feedback Processes in a Dynamical System: Lorenz Model Case Study

A new approach is presented for the analysis of feedback processes in a nonlinear dynamical system by observing its variations. The new methodology consists of statistical estimates of the sensitivities between all pairs of variables in the system based on a neural network modeling of the dynamical system. The model can then be used to estimate the instantaneous, multivariate and nonlinear sensitivities, which are shown to be essential for the analysis of the feedbacks processes involved in the dynamical system. The method is described and tested on synthetic data from the low-order Lorenz circulation model where the correct sensitivities can be evaluated analytically.

Aires, Filipe↗

Leveraging Machine Learning Capabilities for the Characterization of Irradiated Uranium: A Case Study of Analysis Methods for Nuclear Safeguards and Nuclear Forensics

Nondestructively determining the initial enrichment of irradiated uranium is a complex and laborious multivariable problem due to the presence of fission products. This work demonstrates the capabilities of machine learning to analyze gamma-ray spectral data to determine initial enrichment without knowledge of the decay time of the sample. The approach developed is agnostic to the particular scenario and is applicable to a wide variety of applications in nuclear forensics and nuclear safeguards. We irradiated 5 mg uranium standard reference materials at discrete enrichment values ranging from 0.02% to 97% 235 U (weight percent) in UT Austin’s Nuclear Engineering Teaching Laboratory TRIGA Mark II 1.1 MW research reactor, allowed each to decay for 8 hours, and then measured each sample via gamma-ray spectrometry for 50 hours post-irradiation yielding 1,400 individual gamma-ray spectra discretized into 8,192 energy bins. We then trained decision trees models to analyze individual gamma-ray spectra and estimate the associated initial enrichment without knowledge of the time since end of irradiation. We evaluated the performance of the models with a reserved test set not used for training or calibrating the model. A decision tree model constructed with this procedure achieved a mean absolute error in initial enrichment determination of 2.3% (weight percent 235 U). Next, we implemented a principal component analysis pre-processing routine of the gamma-ray spectrometry data to reduce the dimensionality of the dataset from 8,192 channels in the spectrum to 10 principal components while retaining over 99% of the inherent variance in the data. Decision tree models constructed with these data demonstrated decreased mean absolute error in enrichment determination, reduced computational time, and decreased complexity. A single decision tree model constructed with this procedure achieved a mean absolute error in initial enrichment determination of 0.05% (weight percent 235 U). Furthermore, we analyzed these models with learning curves to ensure that overfitting did not occur. The capabilities provided by these models can be naturally extended to other application-focused measurements in the fields of nuclear safeguards, nuclear forensics, and nuclear non-proliferation.

Drescher, Adam↗

Updates to Relevance Vector Machine: Multiclass Classification, Variable Selection, and Proof-of-Concept Application to Safeguards Fresh Fuel Verification using List-Mode Neutron Collar Data

To expand the capabilities of safeguards authorities to verify the integrity of fresh fuel assemblies, Oak Ridge National Laboratory has retrofit the existing electronics of the JCC-71 uranium neutron coincidence collar, which contains 18 3 He neutron detectors and an external 241 AmLi(α, n) neutron interrogation source arranged to surround a fresh nuclear fuel assembly. The new electronics system allows analysts to record list-mode neutron multiplicity data in addition to the singles and doubles rates that are currently measured. Based on previous proof-of-concept research, analysis of these new data will identify off-normal fuel configurations in an assembly and characterize or localize the specific partial fuel defects. The purpose of this report it to document the analysis algorithm development and then to demonstrate its capability for the safeguards verification of fresh fuel assemblies using list mode neutron collar data. To analyze the complex list-mode data collected with the upgraded uranium neutron collar, multivariate classification algorithms are being developed using a novel classification method, the relevance vector machine. This approach may be applied to multiclass problems to estimate the probability that test data belongs to one of many possible classes of data. In addition, our method identifies the most useful variables/channels for making predictions, which illuminates the basis for the model’s predictions, and this interpretability is largely unique among data analytics methods. Variable selection occurs during model training and parameter tuning and does not need any external hyperparameter tuning routines. Finally, we apply the modified relevance vector machine to a simulated dataset of list-mode neutron collar data generated with the radiation transport code MCNP. The method can correctly identify off-normal fuel configurations, categorize the data according to four fuel defect scenarios, and rank the channels in the data according to prediction utility. For nuclear safeguards applications, it is concluded that this method has the potential to increase the sensitivity and reliability to detect missing fuel rods from a standard 17 x 17 Pressurized Water Reactor (PWR) fresh fuel assembly. Within this analysis, “off-normal” (i.e., missing fuel rods) were correctly classified in 17 simulated test scenarios with one quarter (25%) of the fresh fuel rods missing using a training data set of 58 simulated measurements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Multivariate Analysis, Retrieval, and Storage System (MARS). Volume 6: MARS System - A Sample Problem (Gross Weight of Subsonic Transports)

The Mars system is a tool for rapid prediction of aircraft or engine characteristics based on correlation-regression analysis of past designs stored in the data bases. An example of output obtained from the MARS system, which involves derivation of an expression for gross weight of subsonic transport aircraft in terms of nine independent variables is given. The need is illustrated for careful selection of correlation variables and for continual review of the resulting estimation equations. For Vol. 1, see N76-10089.

Hague, D. S.↗

Measurement of the top quark pole mass using $ \textrm{t}\overline{\textrm{t}} $+jet events in the dilepton final state in proton-proton collisions at $ \sqrt{s} $ = 13 TeV

A measurement of the top quark pole mass $\mathcal{m}^\text{pole}_\text{t}$ in events where a top quark-antiquark pair ($\text{t}\bar{\text{t}}$) is produced in association with at least one additional jet ($\text{t}\bar{\text{t}}$ +jet) is presented. This analysis is performed using proton-proton collision data at $\sqrt{s}$ = 13 TeV collected by the CMS experiment at the CERN LHC, corresponding to a total integrated luminosity of 36.3 fb -1 . Events with two opposite-sign leptons in the final state (e + e – , μ + μ – , e ± μ ∓ ) are analyzed. The reconstruction of the main observable and the event classification are optimized using multivariate analysis techniques based on machine learning. The production cross section is measured as a function of the inverse of the invariant mass of the $\text{t}\bar{\text{t}}$ +jet system at the parton level using a maximum likelihood unfolding. Given a reference parton distribution function (PDF), the top quark pole mass is extracted using the theoretical predictions at next-to-leading order. For the ABMP16NLO PDF, this results in $\mathcal{m}^\text{pole}_\text{t}$ = 172.93 ± 1.36 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multi-Gas Sensors for Enhanced Reliability of SOFC Operation

GE Research, in partnership with SUNY Polytechnic Institute (SUNY Poly), designed, built, and tested gas sensors for in situ monitoring of H 2 and CO anode tail gases produced with on-site steam reforming in solid oxide fuel cell (SOFC) systems. The knowledge of the H 2 /CO ratio of these anode tail gases should allow accurate determination and control of the efficiency of the reforming process in the SOFC system and should deliver a lower operating cost for SOFC customers. The project objectives were to achieve multi-gas monitoring capability with a single multivariable sensor, and to sustain this performance in the presence of gaseous interferences and a potential poison for the sensor. The duration of the project was 24 months with the project structure that included three technical tasks such as (1) development of design rules of photonic nanostructures for H 2 and CO gas detection, (2) laboratory validation of photonic nanostructures for selective H 2 and CO gas detection, and (3) validation of photonic nanostructures for initial stability and poison-resistance against H 2 S. To build multivariable sensors for selective H 2 and CO gas detection in the presence of interferences, we expanded our earlier knowledge of multi-gas sensors into new fabrication and functionalization methodologies as well as into new methodologies for the spectral data analysis of multi-gas responses. We have advanced our design rules of the three-dimensional (3D) photonic nanostructures that allowed detection of H 2 and CO at high temperatures as individual gases and as their mixtures and rejection of interferences such as CO 2 , H 2 O, CH 4 , and other hydrocarbons for SOFC applications. Our advanced design rules should be attractive for building the new generation of cost-effective industrial sensors. Stability and poison resistance of our 3D photonic nanostructures was tested in the laboratory conditions. While initially we utilized conventional machine learning data analysis tools, we have found that they were unable to correct for the sensor drift. Thus, we have implemented new methods of machine learning for the analysis of our spectral data. These learnings pave the way to move the future studies into advanced testing of effects of interferences, aging and field tests. In future, our work will continue to advance our sensing designs to operate in conditions with known and unknown interferences by implementing nanostructures with enhanced spectral diversity of responses to gaseous species of interest and interferences. Our systematic reduction of technical risks in this completed project and in future studies will ensure transition of this sensing technology to commercialization.

03 NATURAL GAS↗

Quantitative Comparison of the Variability in Observed and Simulated Shortwave Reflectance

The Climate Absolute Radiance and Refractivity Observatory (CLARREO) is a climate observation system that has been designed to monitor the Earth's climate with unprecedented absolute radiometric accuracy and SI traceability. Climate Observation System Simulation Experiments (OSSEs) have been generated to simulate CLARREO hyperspectral shortwave imager measurements to help define the measurement characteristics needed for CLARREO to achieve its objectives. To evaluate how well the OSSE-simulated reflectance spectra reproduce the Earth s climate variability at the beginning of the 21st century, we compared the variability of the OSSE reflectance spectra to that of the reflectance spectra measured by the Scanning Imaging Absorption Spectrometer for Atmospheric Cartography (SCIAMACHY). Principal component analysis (PCA) is a multivariate decomposition technique used to represent and study the variability of hyperspectral radiation measurements. Using PCA, between 99.7%and 99.9%of the total variance the OSSE and SCIAMACHY data sets can be explained by subspaces defined by six principal components (PCs). To quantify how much information is shared between the simulated and observed data sets, we spectrally decomposed the intersection of the two data set subspaces. The results from four cases in 2004 showed that the two data sets share eight (January and October) and seven (April and July) dimensions, which correspond to about 99.9% of the total SCIAMACHY variance for each month. The spectral nature of these shared spaces, understood by examining the transformed eigenvectors calculated from the subspace intersections, exhibit similar physical characteristics to the original PCs calculated from each data set, such as water vapor absorption, vegetation reflectance, and cloud reflectance.

Roberts, Yolanda, L.↗

Numerical Reanalyses as a Gateway to Arctic Synthesis

Reanalyses are regularly gridded, retrospective depictions of the physical earth system, which are produced through the correction of a short-term forecast to available observations. In the Arctic, reanalyses are particularly well suited to marshal the sparse observing network to provide a plausible, multivariate representation of conditions. Atmospheric reanalyses such as MERRA-2 (NASA Modern-Era Retrospective analysis for Research and Applications, version 2) and ocean reanalyses such as SODA3 (Univ. Maryland Simple Ocean Data Assimilation version 3) are widely used in Arctic research for diagnostic studies of circulation, model evaluation, and as boundary conditions for a variety of process models. Here, we provide examples that illustrate the utility of reanalyses for providing information on the spatial and temporal scales of recent, rapid changes in the Arctic. Recent trends in Arctic surface temperatures, surface melt over Greenland and Arctic glaciers, and evolving freshwater conditions in the Arctic Ocean are examples where reanalyses can provide information that cannot easily be obtained via other means. These examples provide information on the scale, magnitude, and the uncertainty of recent Arctic change and provide a context for future scenarios. We further quantify uncertainties in key reanalyses variables and approaches for addressing these issues.

Arctic↗