Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Leveraging Machine Learning Capabilities for the Characterization of Irradiated Uranium: A Case Study of Analysis Methods for Nuclear Safeguards and Nuclear Forensics

Nondestructively determining the initial enrichment of irradiated uranium is a complex and laborious multivariable problem due to the presence of fission products. This work demonstrates the capabilities of machine learning to analyze gamma-ray spectral data to determine initial enrichment without knowledge of the decay time of the sample. The approach developed is agnostic to the particular scenario and is applicable to a wide variety of applications in nuclear forensics and nuclear safeguards. We irradiated 5 mg uranium standard reference materials at discrete enrichment values ranging from 0.02% to 97% 235 U (weight percent) in UT Austin’s Nuclear Engineering Teaching Laboratory TRIGA Mark II 1.1 MW research reactor, allowed each to decay for 8 hours, and then measured each sample via gamma-ray spectrometry for 50 hours post-irradiation yielding 1,400 individual gamma-ray spectra discretized into 8,192 energy bins. We then trained decision trees models to analyze individual gamma-ray spectra and estimate the associated initial enrichment without knowledge of the time since end of irradiation. We evaluated the performance of the models with a reserved test set not used for training or calibrating the model. A decision tree model constructed with this procedure achieved a mean absolute error in initial enrichment determination of 2.3% (weight percent 235 U). Next, we implemented a principal component analysis pre-processing routine of the gamma-ray spectrometry data to reduce the dimensionality of the dataset from 8,192 channels in the spectrum to 10 principal components while retaining over 99% of the inherent variance in the data. Decision tree models constructed with these data demonstrated decreased mean absolute error in enrichment determination, reduced computational time, and decreased complexity. A single decision tree model constructed with this procedure achieved a mean absolute error in initial enrichment determination of 0.05% (weight percent 235 U). Furthermore, we analyzed these models with learning curves to ensure that overfitting did not occur. The capabilities provided by these models can be naturally extended to other application-focused measurements in the fields of nuclear safeguards, nuclear forensics, and nuclear non-proliferation.

Drescher, Adam↗

Time Series Classification for Locating Forced Oscillation Sources

Here, this article presents a machine learning based time-series classification method for using synchrophasor measurements to locate the source of forced oscillation (FO) for fast disturbance removal. First, multivariate time series (MTS) matrices are constructed by the most informative measurements selected by sequential feature selection from each power plant. Then, the Mahalanobis matrix is trained such that the Mahalanobis distance between the MTSs from the same class (i.e., with the same FO source location) are minimized and from different classes (i.e., with different FO source locations) are maximized. This allows MTSs to be classified by classifiers with class membership corresponding to the location of each FO source. To meet the runtime requirements of online matching, class templates are constructed to reduce data size and improve matching efficiency. To account for uncertainty in identifying the exact beginning of an FO event, dynamic time warping is used to align the out-of-sync MTSs. IEEE 39bus and WECC 179bus systems are used for algorithm development and validation. Simulation results demonstrate that the algorithm meets online operation runtime requirement with high accuracy using misaligned data sets.

42 ENGINEERING↗

High‐Resolution National‐Scale Water Modeling Is Enhanced by Multiscale Differentiable Physics‐Informed Machine Learning

Abstract The National Water Model (NWM) is a key tool for flood forecasting, planning, and water management. Key challenges facing the NWM include calibration and parameter regionalization when confronted with big data. We present two novel versions of high‐resolution (∼37 km 2 ) differentiable models (a type of hybrid model): one with implicit, unit‐hydrograph‐style routing and another with explicit Muskingum‐Cunge routing in the river network. The former predicts streamflow at basin outlets whereas the latter presents a discretized product that seamlessly covers rivers in the conterminous United States (CONUS). Both versions use neural networks to provide a multiscale parameterization and process‐based equations to provide a structural backbone, which were trained simultaneously (“end‐to‐end”) on 2,807 basins across the CONUS and evaluated on 4,997 basins. Both versions show great potential to elevate future NWM performance for extensively calibrated as well as ungauged sites: the median daily Nash‐Sutcliffe efficiency of all 4,997 basins is improved to around 0.68 from 0.48 of NWM3.0. As they resolve spatial heterogeneity, both versions greatly improved simulations in the western CONUS and also in the Prairie Pothole Region, a long‐standing modeling challenge. The Muskingum‐Cunge version further improved performance for basins >10,000 km 2 . Overall, our results show how neural‐network‐based parameterizations can improve NWM performance for providing operational flood predictions while maintaining interpretability and multivariate outputs. The modeling system supports the Basic Model Interface (BMI), which allows seamless integration with the next‐generation NWM. We also provide a CONUS‐scale hydrologic data set for further evaluation and use.

Song, Yalan [Civil and Environmental Engineering T↗

The Chemistry Graduate Student Experience: Findings from an ACS Survey

Graduate training is a key element in producing a scientific workforce that reflects the nation’s diversity. This paper examines data from a 2013 American Chemical Society (ACS) survey of 2,544 chemistry masters and doctoral students and reveals barriers to reaching this goal. Multivariate statistical analyses indicate that women reported significantly less supportive relationships with advisors. Women were less likely to plan to finish their degrees, and for PhD students, the discrepancy was larger for students at the start of their graduate program. Women were also less likely to pursue the next level of training, and the gender difference related to postdoctoral plans was greater for those who identified with a racial-ethnic group traditionally underrepresented in chemistry (underrepresented minority, URM). URM students who were beyond the first year of their graduate program reported significantly less supportive relationships with peers. They were also less likely to have funding sufficient to meet their needs and more often used personal resources including loans. Despite these difficulties, URM students were more likely to definitely plan to finish their degrees, and men who identified as URM were more likely to plan to pursue postdoctoral work. Independent of gender and identification as URMs, students in more highly ranked schools reported less advisor support. Extensive open-ended comments indicated that large proportions of the students desired more attention and meaningful feedback from advisors and changes within their programs to promote support for students and advisor accountability. Suggestions for future research are given, and a companion commentary discusses needed directions for change.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MSW Variability Mapping and Conversion to Biofuel

MSW (Municipal Solid Waste) is a form of biomass which consists of categorized components of waste/trash. The general categories are paper, yard trash, construction & debris, appliances, tires, glass, metals, aluminum & steel cans, plastics, organics, inorganics, and HHW (Household Hazardous Waste). This project focuses on the factors within a region or population that contribute to variability in the composition of MSW and in turn MSW’s convertibility to biofuel. A list of contributors was determined (Social Vulnerability Index, Access to Public Transportation, Racial Distribution, GDP, Personal Income) and then JMP was used to perform a Multivariate analysis to determine correlations and a Partial Least-Squares regression to determine Variable Importance Plots for each MSW category. In addition to data analysis, the convertibility of MSW to biofuel was studied via microwave pyrolysis system in order to separate and characterize the various gaseous and bio-oil products.

09 BIOMASS FUELS↗

A hidden Markov model for population-level cervical cancer screening data

The Cancer Registry of Norway has been administrating a national cervical cancer screening program since 1992 by coordinating triennial cytology exam screenings for the female population between 25 and 69 years of age. Up to 80% of cancers are prevented through mass screening, but this comes at the expense of considerable screening activity and leads to overtreatment of clinically asymptomatic precancers. Here, we present a continuous-time, time-inhomogeneous hidden Markov model which was developed to understand the screening process and cervical cancer carcinogenesis in detail. By leveraging 1.7 million individual's multivariate time-series of medical exams performed over a 25-year period, we simultaneously estimate all model parameters. We show that an age-dependent model reflects the Norwegian screening program by comparing empirical survival curves from observed registry data and data simulated from the proposed model. The model can be generalized to include more detailed individual-level covariates as well as new types of screening exams. By utilizing individual screening histories and covariate data, the proposed model shows potential for improving strategies for cancer screening programs by personalizing recommended screening intervals.

59 BASIC BIOLOGICAL SCIENCES↗

Switchgrass sward establishment selection is consistent across multiple environments and fertilization levels

Strong selection can occur during switchgrass sward establishment. Differences in establishment selection due to environment or management could provide information on genotype-by-environment variation and could influence strategies for breeding perennial grasses. Leaf samples were collected before sward establishment and from 3-year-old swards for two breeding groups (lowland and hybrid) at three locations. Within two locations, samples were collected from paired fertilized (112 kg N ha –1 ) and unfertilized plots. Allele frequencies from pooled DNA samples were studied through multivariate analysis of variance, genomewide trait predictions (heading date and winter survivorship), and genomically estimated breeding values (GEBVs) for individual sward survival within an independent data set. This study found only minor variations in selection due to location or management. Predicted heading dates of the hybrid population had significant changes due to fertilization and location. There were strong correlations among sward establishment survival GEBVs between growing environments (hybrid r = 0.77; gulf r = 0.97). Interestingly, this study found a small number of genotypes that were over-represented in established swards across all growing environments. This study reinforces a prior report of selection during sward establishment and indicates that only a small degree of establishment selection is location-specific within these diverse growing conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Perspectives on Polyolefin Catalysis in Microfluidics for High-Throughput Screening: A Minireview

Polyolefins are the largest produced plastics in the world which traditionally employ continuous stirred tank reactors and fluidized bed reactors for commercial production. The operating condition, reaction kinetics, and molecular interactions inside the reactor strongly affect the polyolefin properties, which require stringent process control in conventional procedures. Understanding the catalytic pathway, behavior of polymer particles and effect of reactor conditions are essential for designing specific polymer properties, namely the molecular weight, chain length, polydispersity, etc. Microfluidics can play a significant role in designing polymers tailored to the user needs. Smaller channel dimensions help obtain uniform reaction conditions over the length of the microfluidic reactor in a controlled environment. With real-time monitoring techniques in microfluidics, even single particle growth of polymer can be studied to understand the parameters affecting the polymer properties. High throughput microfluidics can help catalyst screening in a short duration with less consumption of reagents generating less waste. When supplemented with efficient machine learning algorithms, automated high throughput microfluidics has the potential to rapidly optimize the process and develop new knowledge even with a limited data set. When trained on data sets generated using microfluidic experiments that are designed efficiently with working knowledge of the process, machine learning algorithms can provide the relationship between the multivariable parameters space and polymer properties, which is not possible with the traditional statistical methods and interpolation techniques. Here, the rise in the utilization of microfluidics, with the advancement of machine learning algorithms, for polyolefin catalysis, highlights the importance of microfluidics for catalyst discovery, parameter optimization, and understanding reaction pathway for producing polymers with specific properties for specialized applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Southern Photometric Quasar Catalog from the Dark Energy Survey Data Release 2

We present a catalog of 1.4 million photometrically selected quasar candidates in the southern hemisphere over the ~5000 deg 2 Dark Energy Survey (DES) wide survey area. We combine optical photometry from the DES second data release (DR2) with available near-infrared (NIR) and the all-sky unWISE mid-infrared photometry in the selection. We build models of quasars, galaxies, and stars with multivariate skew-t distributions in the multidimensional space of relative fluxes as functions of redshift (or color for stars) and magnitude. Our selection algorithm assigns probabilities for quasars, galaxies, and stars and simultaneously calculates photometric redshifts (photo-z) for quasar and galaxy candidates. Benchmarking on spectroscopically confirmed objects, we successfully classify (with photometry) 94.7% of quasars, 99.3% of galaxies, and 96.3% of stars when all IR bands (NIR YJHK and WISE W1W2) are available. The classification and photo-z regression success rates decrease when fewer bands are available. Our quasar (galaxy) photo-z quality, defined as the fraction of objects with the difference between the photo-z z p and the spectroscopic redshift z s , |Δz| ≡ |z s - z p |/(1 + z s ) ≤ 0.1, is 92.2% (98.1%) when all IR bands are available, decreasing to 72.2% (90.0%) using optical DES data only. Our photometric quasar catalog achieves an estimated completeness of 89% and purity of 79% at r < 21.5 (0.68 million quasar candidates), with reduced completeness and purity at 21.5 < r ≲ 24. Among the 1.4 million quasar candidates, 87,857 have existing spectra, and 84,978 (96.7%) of them are spectroscopically confirmed quasars. Finally, we provide quasar, galaxy, and star probabilities for all (0.69 billion) photometric sources in the DES DR2 coadded photometric catalog.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimizing high energy density sulfur cathodes: A multivariate approach to electrode formulation and processing

Lithium-sulfur (Li-S) batteries involve complex solid-liquid-solid phase transformations during both discharging and charging processes, where cathode materials, formulation, and structure play a crucial role. Here, a design of experiments (DoE) methodology and an empirical model are developed to systematically explore the interactions and trade-offs among cathode factors and process variables, and to obtain generalizable effects estimates for the multivariate system. Compared to the conventional one-factor-at-a-time (OFAT) approach, this work demonstrates advantages in both efficiency and accuracy by allowing the data to guide future research and decisions. Further, an optimized cathode formulation and processing parameters are predicted and validated experimentally, achieving over 1000 mAh g -1 in discharge capacity and improved cycling under practical lean electrolyte (4 µL mg -1 S) and high S-loading cathodes (>4 mg cm -2 ) conditions. The optimized cathode was scaled up and assembled into Li-S pouch cells, achieving 316 Wh kg -1 in cell-level energy, proving that the comprehensive and rigorous framework for optimizing complex systems with DoE leads to improved performance in a practical pouch cell system.

25 ENERGY STORAGE↗

A Transported Livengood–Wu Integral Model for Knock Prediction in Computational Fluid Dynamics Simulation

This work describes the development of a transported Livengood–Wu (L–W) integral model for computational fluid dynamics (CFD) simulation to predict autoignition and engine knock tendency. The currently employed L–W integral model considers both single-stage and two-stage ignition processes, thus can be generally applied to different fuels such as paraffin, olefin, aromatics, and alcohol. The model implementation is first validated in simulations of homogeneous charge compression ignition (HCCI) combustion for three different fuels, showing good accuracy in prediction of autoignition timing for fuels with either single-stage or two-stage ignition characteristics. Then, the L–W integral model is coupled with G-equation model to indicate end-gas autoignition and knock tendency in CFD simulations of a direct-injection spark-ignition engine. This modeling approach is about 10 times more efficient than the ones that based on detailed chemistry calculation and pressure oscillation analysis. Two fuels with same Research Octane Number (RON) but different octane sensitivity are studied, namely, Co-Optima alkylate and Co-Optima E30. Feed-forward neural network model in conjunction with multivariable minimization technique is used to generate fuel surrogates with targets of matched RON, octane sensitivity, and ethanol content. The CFD model is validated against experimental data in terms of pressure traces and heat release rate for both fuels under a wide range of operating conditions. The knock tendency—indicated by the fuel energy contained in the autoignited region—of the two fuels at different load conditions correlates well with the experimental results and the fuel octane sensitivity, implying the current knock modeling approach can capture the octane sensitivity effect and can be applied to further investigation on composition of octane sensitivity.

33 ADVANCED PROPULSION SYSTEMS↗

Learning Planar Ising Models Software

Learning Planar Ising Models is a software package written in Matlab for learning relationships among variable in a dataset using graphical models. The software package implements a generally-applicable algorithm for learning planar Ising models from any multivariate dataset. The code provides an algorithm for learning the best planar Ising model to approximate an arbitrary collection of binary random variables (possibly from sample data). Given the set of all pairwise correlations among variables, we select a planar graph and optimal planar Ising model defined on this graph to best approximate that set of correlations. The software includes demonstrations of the algorithm in simulations and for applications on publicly available datasets. Details of the algorithm, demonstration simulations, and applications are given in Johnson, et al; 2016. Reference: Johnson, J. K., Oyen, D., Chertkov, M., and Netrapalli, P. (2016). Learning planar Ising models. Journal of Machine Learning Research.

Oyen, Diane↗

Factors Influencing Willingness to Pool in Ride-Hailing Trips

In the past decade, transportation network companies (TNCs) such as Uber, Lyft, and Via have established themselves as a viable transportation alternative to other modes. However, the popularity of these services has come with a fair share of criticism for their negative externalities such as increasing vehicle miles traveled and congestion in cities. Pooled ride-hailing trips, in which all or a part of two individual (or group) trips are combined in and served by a single vehicle, have the potential to reduce these externalities. Pooling of rides is an effective solution to reduce congestion and travel cost, but pooled rides still represent a small percentage of the total trips served (and miles driven) by TNCs relative to single-occupancy (and without customer) vehicle miles. Both TNCs and cities alike will benefit from understanding what factors encourage or deter pooling a ride-hailing trip. In this study, newly available Chicago transportation network provider data were explored to identify the extent to which different socioeconomic, spatiotemporal, and trip characteristics affect willingness to pool (WTP) in ride-hailing trips. Furthermore, multivariate linear regression and machine-learning models were employed to understand and predict WTP based on location, time, and trip factors. The results show intuitive trends, with income level at drop-off and pickup locations and airport trips as the most important predictors of WTP. Results from this study can help TNCs and cities devise strategies that increase pooled ride-hailing, thereby reducing adverse transportation and energy impacts from ride-hailing modes.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Reducing systematic bias in machine learning applications to J/ψ signal extraction in high-energy nuclear physics

Machine learning techniques are increasingly used in high-energy nuclear physics because they can exploit multivariate correlations more efficiently than conventional cut-based analyses. A central challenge is the construction of training samples that faithfully reproduce the detector response observed in data. Signal samples are usually derived from detector simulations; therefore, mismatches between simulation and data can degrade classifier performance and introduce systematic biases. This work presents two practical correction procedures, namely cumulative distribution function (CDF) mapping and a shift-and-scale transformation, to align simulated signal features with those measured in data. Their performance is demonstrated with $J$/$\psi$ yield measurements in $\sqrt{s_{nn}}$ = 200 GeV Ru+Ru and Zr+Zr collisions recorded by STAR. A set of self-consistency tests shows that these procedures substantially suppress the systematic bias associated with data-simulation discrepancies in machine-learning-based signal extraction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Utilizing the Dynamic Networks Data Processing and Analysis Experiment (DNE18) to Establish Methodologies for the Comparison of Automatic Infrasonic Signal Detectors

The Dynamic Networks Experiment 2018 (DNE18) was a collaborative effort between Los Alamos National Laboratory (LANL), Sandia National Laboratories (SNL), Lawrence Livermore National Laboratory (LLNL) and Pacific Northwest National Laboratory (PNNL) designed to evaluate methodologies for multi-modal data ingestion and processing. One component of this virtual experiment was a quantitative assessment of current capabilities for infrasound data processing, beginning with the establishment of a baseline for infrasound signal detection. To produce such baselines, SNL and LANL exploited a common dataset of infrasound data recorded across a regional network in Utah from December 2010 through February 2011. We utilize two automated signal detectors, the Adaptive F-Detector (AFD) and the Multivariate Adaptive Learning Detector (MALD) to produce automated signal detection catalogs and an analyst-produced catalog. Comparisons indicate that automatic detectors may be able to identify small amplitude, low SNR events that cannot be identified by analyst review. We document detector performance in terms of precision and recall, demonstrating that the AFD is more precise, but the MALD has higher recall. We use a synthetic dataset of signals embedded in pink noise in order to highlight shortcomings in assessing detection algorithms for low signal to noise ratio signals which are commonly of interest to the nuclear monitoring community. For comparisons utilizing the synthetic dataset, the AFD has higher recall while precision is equal for both detectors. These results indicate that both detectors perform well across a variety of background noise environments; however, both detectors fail to identify repetitive, short duration signals arriving from similar backazimuths. These failures represent specific scenarios that could be targeted for further detector development.

97 MATHEMATICS AND COMPUTING↗

A Data-Driven Nonparametric Approach for Probabilistic Load-Margin Assessment Considering Wind Power Penetration

A modern power system is characterized by an increasing penetration of wind power, which results in large uncertainties in its states. These uncertainties must be quantified properly; otherwise, the system security may be threatened. Facing this challenge, here we propose a cost-effective, data-driven approach to assessing a power system's load margin probabilistically. Using actual wind data, a kernel density estimator is applied to infer the nonparametric wind speed distributions, which are further merged into the framework of a vine copula. The latter enables us to simulate complex multivariate and highly dependent model inputs with a variety of bivariate copulae that precisely represent the tail dependence in the correlated samples. Furthermore, to reduce the prohibitive computational time of traditional Monte-Carlo simulations that process a large amount of samples, we propose to use a nonparametric, Gaussian-process-emulator-based reduced-order model to replace the original complicated continuation power-flow model through a Bayesian-learning framework. To accelerate the convergence rate of this Bayesian algorithm, a truncated polynomial chaos surrogate, which serves as a highly efficient, parametric Bayesian prior, is developed. This emulator allows us to execute the time-consuming continuation power-flow solver at the sampled values with a negligible computational cost. Results of simulations that are performed on several test systems reveal the impressive performance of the proposed method in the probabilistic load-margin assessment.

17 WIND ENERGY↗