Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

PM 2.5 Is Insufficient to Explain Personal PAH Exposure

To understand how chemical exposure can impact health, researchers need tools that capture the complexities of personal chemical exposure. In practice, fine particulate matter (PM 2.5 ) air quality index (AQI) data from outdoor stationary monitors and Hazard Mapping System (HMS) smoke density data from satellites are often used as proxies for personal chemical exposure, but do not capture total chemical exposure. Silicone wristbands can quantify more individualized exposure data than stationary air monitors or smoke satellites. However, it is not understood how these proxy measurements compare to chemical data measured from wristbands. In this study, participants wore daily wristbands, carried a phone that recorded locations, and answered daily questionnaires for a 7-day period in multiple seasons. We gathered publicly available daily PM 2.5 AQI data and HMS data. We analyzed wristbands for 94 organic chemicals, including 53 polycyclic aromatic hydrocarbons. Wristband chemical detections and concentrations, behavioral variables (e.g., time spent indoors), and environmental conditions (e.g., PM 2.5 AQI) significantly differed between seasons. Machine learning models were fit to predict personal chemical exposure using PM 2.5 AQI only, HMS only, and a multivariate feature set including PM 2.5 AQI, HMS, and other environmental and behavioral information. On average, the multivariate models increased predictive accuracy by approximately 70% compared to either the AQI model or the HMS model for all chemicals modeled. This study provides evidence that PM 2.5 AQI data alone or HMS data alone is insufficient to explain personal chemical exposures. Our results identify additional key predictors of personal chemical exposure.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

PyKrev: A Python Library for the Analysis of Complex Mixture FT-MS Data

In this study, we present PyKrev, a Python library for the analysis of complex mixture Fourier transform mass spectrometry (FT-MS) data. PyKrev is a comprehensive suite of tools for analysis and visualization of FT-MS data after formula assignment has been performed. These comprise formula manipulation and calculation of chemical properties, intersection analysis between multiple lists of formulas, calculation of chemical diversity, assignment of compound classes to formulas, multivariate analysis, and a variety of visualization tools producing van Krevelen diagrams, class histograms, PCA score, and loading plots, biplots, scree plots, and UpSet plots. The library is showcased through analysis of hot water green tea extracts and Scotch whisky FT-ion cyclotron resonance-MS data sets. PyKrev addresses the lack of a single, cohesive toolset for researchers to perform FT-MS analysis in the Python programming environment encompassing the most recent data analysis techniques used in the field.

47 OTHER INSTRUMENTATION↗

Multiphenomenology explosion monitoring (MultiPEM): a general framework for data interpretation and yield estimation

SUMMARY An underground nuclear explosion (UNE) couples mechanical energy into crustal rock, which propagates as seismic and acoustic waves. These different physical phenomena transport, by different pathways, to standoff detectors at varying distances. The transport pathways attenuate the original signal but in different ways. Enabled by correct statistical weighting, signal attenuation models can be used to combine these disparate sensor data to estimate the yield of an UNE. Contemporaneous statistical models, used in yield estimation, can be improved with an advanced partition of error for these physical signal propagation models. We present an advanced multivariate approach to error modelling of multiphenomenology physical signatures. In addition to measurement error, our error model represents physical model biases as random with a physics-based covariance structure. To illustrate this proposed framework, we demonstrate the estimation of explosion yield using openly available seismic and acoustic data from chemical single-point explosions.

Williams, Brian J.↗

Accelerating Multivariate Functional Approximation Computation with Domain Decomposition Techniques⋆

Modeling large datasets through Multivariate Functional Approximations (MFA) provide an elegant way to handle many visualization and scientific analysis workflows. The process necessitates scalable data partitioning methods to compute MFA representations efficiently without compromising the accuracy or continuity of the reconstructed solution. We propose a domain -decomposed method for computing the MFA with B -spline bases, which reduces the total work per task and uses a restricted Additive Schwarz (RAS) method to converge the control point data degrees -of -freedom along subdomain boundaries. We provide an in-depth analysis of the parallel approach with domain decomposition solvers, aiming to minimize local subdomain error residuals and recover high -order continuity at subdomain interfaces with appropriate choices of knot overlaps. The communication cost, determined by the overlap regions in the RAS implementation, is optimized to recover the numerical error profile of the single subdomain case. Our proposed method stands in contrast to previous methods, which typically only recover either C 0 or at best C 1 continuity for arbitrary B -spline degree expansions, or those that require post -processing to blend discontinuities in the reconstructed data. We demonstrate the effectiveness of our approach using analytical and real -world datasets in 1D, 2D, and 3D through both strong and weak scaling studies. The performance results indicate that the overall cost of computing the approximation is directly proportional to the underlying nearest -neighbor communication implementation, and is only weakly dependent on the overlap region size that determines the size of the messages. This finding underscores the efficiency and scalability of our proposed method, making it a promising solution for handling large datasets in scientific workflows.

additive Schwarz solvers↗

Predicting Postoperative Injury and Military Discharge Status After Knee Surgery in the US Army

Background: Researchers have assessed postoperative injury or disability predictors in the military setting but typically focused on 1 type of surgical procedure at a time, used relatively small sample sizes, or investigated mixed cohorts with civilian populations. Purpose: To identify the relationship between baseline variables and injury incidence or military discharge status in US Army soldiers after knee surgery. Study Design: Case-control study; Level of evidence, 3. Methods: Data were obtained from a repository containing personnel, performance, and medical records for all active-duty US Army soldiers. Multivariate logistic regressions were used to estimate the effects of numerous variables on postoperative injury or on medical discharge. Variable selection and model validation were conducted using the k-fold method. Results: A total of 7567 soldiers underwent knee surgery between 2017 and 2019. Meniscal procedures were the most common type of surgery (39%), and approximately 71% of the cohort had a postoperative injury. Significant predictors for sustaining a postoperative injury included having a previous nonknee injury (odds ratio [OR], 1.5), female sex (OR, 1.3), and Black race (OR, 1.2). Within 4 years after surgery, 17% of soldiers were discharged from the military because of knee-related disability. Significant predictors for discharge from duty included enlisted rank (OR, 2.3), recent fitness test failure (OR, 1.9), number of previous knee surgeries (OR, 1.7), and having a previous nonknee injury (OR, 1.6). Conclusion: After knee surgery, nearly three-fourths of the soldiers in this cohort sustained a postoperative injury and almost one-fifth of soldiers were medically discharged from the military within 4 years. This study identified variables that indicate statistically increased risk for these postoperative outcomes and highlighted potentially modifiable factors.

Orthopedics↗

COVID-19 vaccination status, side effects, and perceptions among breast cancer survivors: a cross-sectional study in China

Introduction Breast cancer is the most prevalent malignancy in patients with coronavirus disease 2019 (COVID-19). However, vaccination data of this population are limited. Methods A cross-sectional study of COVID-19 vaccination was conducted in China. Multivariate logistic regression models were used to assess factors associated with COVID-19 vaccination status. Results Of 2,904 participants, 50.2% were vaccinated with acceptable side effects. Most of the participants received inactivated virus vaccines. The most common reason for vaccination was “fear of infection” (56.2%) and “workplace/government requirement” (33.1%). While the most common reason for nonvaccination was “worry that vaccines cause breast cancer progression or interfere with treatment” (72.9%) and “have concerns about side effects or safety” (39.6%). Patients who were employed (odds ratio, OR = 1.783, p = 0.015), had stage I disease at diagnosis (OR = 2.008, p = 0.019), thought vaccines could provide protection (OR = 1.774, p = 0.007), thought COVID-19 vaccines were safe, very safe, not safe, and very unsafe (OR = 2.074, p < 0.001; OR = 4.251, p < 0.001; OR = 2.075, p = 0.011; OR = 5.609, p = 0.003, respectively) were more likely to receive vaccination. Patients who were 1–3 years, 3–5 years, and more than 5 years after surgery (OR = 0.277, p < 0.001; OR = 0.277, p < 0.001, OR = 0.282, p < 0.001, respectively), had a history of food or drug allergies (OR = 0.579, p = 0.001), had recently undergone endocrine therapy (OR = 0.531, p < 0.001) were less likely to receive vaccination. Conclusion COVID-19 vaccination gap exists in breast cancer survivors, which could be filled by raising awareness and increasing confidence in vaccine safety during cancer treatment, particularly for the unemployed individuals.

Xu, Yali↗

STITCHES: a Python package to amalgamate existing Earth system model output into new scenario realizations

Understanding the interaction between humans and the Earth system is a computationally daunting task, with many possible approaches depending on resources available and questions of interest. For example, state-of-the-art impact models require decade-long time series of relatively high frequency, spatially resolved and often multiple variables representing climatic impact-drivers (Ruane et al., 2022). Most commonly these are derived from the outputs of detailed, computationally expensive Earth System Models (ESMs) run according to a standard, limited set of future scenarios, the latest being the SSP-RCPs run under CMIP6/ScenarioMIP (Eyring et al., 2016; O’Neill et al., 2016). At the time of writing, O’Neill et al. (2016) has been cited more than 1750 times and Eyring et al. (2016) more than 5000 times, highlighting the broad, general applications of this data. Often, however, impact modeling seeks to explore new scenarios that were not part of the ScenarioMIP protocol, and/or needs a larger set of initial condition ensemble members than are typically available to quantify the effects of ESM internal variability. In addition, the recognition that the human and Earth systems are fundamentally intertwined, and may feature potentially significant feedback loops, is making integrated, simultaneous modeling of the coupled human-Earth system increasingly necessary, if computationally challenging with most existing tools (Thornton et al., 2017). For impact modelers, climate model emulators can be the answer to meet both the needs of: 1) creating realizations for novel scenarios and 2) achieving a simplified, computationally tractable representation of ESM behavior in a coupled human-Earth system modeling framework. We proposed a new, comprehensive approach to such emulation of gridded, multivariate ESM outputs for novel scenarios without the computational cost of a full ESM, STITCHES (Tebaldi et al., 2022). The approach outlined in Tebaldi et al. (2022) should be extensible to future CMIP eras, although the STITCHES software at present is strictly focused on CMIP6/ScenarioMIP data hosted on Pangeo (https://gallery.pangeo.io/repos/pangeo-gallery/cmip6/). The corresponding STITCHES Python package uses existing archives of ESMs’ scenario experiments from CMIP6/ScenarioMIP to construct gridded, multivariate realizations of new scenarios provided by reduced complexity climate models (Hartin et al., 2015; Meinshausen et al., 2011; Smith et al., 2018), or to enrich existing initial condition ensembles. Its output provides the same characteristics as the emulated ESM output: multivariate (spanning potentially all variables that the ESM has saved), spatially resolved (down to the native grid of the ESM), and preserving the same high frequency as the original data. A new realization of multiple variables can be generated on the order of minutes with STITCHES, rather than the hours or sometimes days that ESMs require.

97 MATHEMATICS AND COMPUTING↗

Dietary B group vitamin intake and the bladder cancer risk: a pooled analysis of prospective cohort studies

Abstract Purpose Diet may play an essential role in the aetiology of bladder cancer (BC). The B group complex vitamins involve diverse biological functions that could be influential in cancer prevention. The aim of the present study was to investigate the association between various components of the B group vitamin complex and BC risk. Methods Dietary data were pooled from four cohort studies. Food item intake was converted to daily intakes of B group vitamins and pooled multivariate hazard ratios (HRs), with corresponding 95% confidence intervals (CIs), were obtained using Cox-regression models. Dose–response relationships were examined using a nonparametric test for trend. Results In total, 2915 BC cases and 530,012 non-cases were included in the analyses. The present study showed an increased BC risk for moderate intake of vitamin B1 (HR B1 : 1.13, 95% CI: 1.00–1.20). In men, moderate intake of the vitamins B1, B2, energy-related vitamins and high intake of vitamin B1 were associated with an increased BC risk (HR (95% CI): 1.13 (1.02–1.26), 1.14 (1.02–1.26), 1.13 (1.02–1.26; 1.13 (1.02–1.26), respectively). In women, high intake of all vitamins and vitamin combinations, except for the entire complex, showed an inverse association (HR (95% CI): 0.80 (0.67–0.97), 0.83 (0.70–1.00); 0.77 (0.63–0.93), 0.73 (0.61–0.88), 0.82 (0.68–0.99), 0.79 (0.66–0.95), 0.80 (0.66–0.96), 0.74 (0.62–0.89), 0.76 (0.63–0.92), respectively). Dose–response analyses showed an increased BC risk for higher intake of vitamin B1 and B12. Conclusion Our findings highlight the importance of future research on the food sources of B group vitamins in the context of the overall and sex-stratified diet.

60 APPLIED LIFE SCIENCES↗

Search for top squarks in the four-body decay mode with single lepton final states in proton-proton collisions at $ \sqrt{s} $ = 13 TeV

A search for the pair production of the lightest supersymmetric partner of the top quark, the top squark ($ {\overset{\sim }{\textrm{t}}}_1 $), is presented. The search targets the four-body decay of the $ {\overset{\sim }{\textrm{t}}}_1 $, which is preferred when the mass difference between the top squark and the lightest supersymmetric particle is smaller than the mass of the W boson. This decay mode consists of a bottom quark, two other fermions, and the lightest neutralino ($ {\overset{\sim }{\chi}}_1^0 $), which is assumed to be the lightest supersymmetric particle. The data correspond to an integrated luminosity of 138 fb$^{−1}$ of proton-proton collisions at a center-of-mass energy of 13 TeV collected by the CMS experiment at the CERN LHC. Events are selected using the presence of a high-momentum jet, an electron or muon with low transverse momentum, and a significant missing transverse momentum. The signal is selected based on a multivariate approach that is optimized for the difference between m($ {\overset{\sim }{\textrm{t}}}_1 $) and m($ {\overset{\sim }{\chi}}_1^0 $). The contribution from leading background processes is estimated from data. No significant excess is observed above the expectation from standard model processes. The results of this search exclude top squarks at 95% confidence level for masses up to 480 and 700 GeV for m($ {\overset{\sim }{\textrm{t}}}_1 $) − m($ {\overset{\sim }{\chi}}_1^0 $) = 10 and 80 GeV, respectively.[graphic not available: see fulltext]

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Statistical Interpolation Code for Ocean Analysis and Forecasting

Abstract We present a data assimilation package for use with ocean circulation models in analysis, forecasting, and system evaluation applications. The basic functionality of the package is centered on a multivariate linear statistical estimation for a given predicted/background ocean state, observations, and error statistics. Novel features of the package include support for multiple covariance models, and the solution of the least squares normal equations either using the covariance matrix or its inverse—the information matrix. The main focus of this paper, however, is on the solution of the analysis equations using the information matrix, which offers several advantages for solving large problems efficiently. Details of the parameterization of the inverse covariance using Markov random fields are provided and its relationship to finite-difference discretizations of diffusion equations are pointed out. The package can assimilate a variety of observation types from both remote sensing and in situ platforms. The performance of the data assimilation methodology implemented in the package is demonstrated with a yearlong global ocean hindcast with a 1/4° ocean model. The code is implemented in modern Fortran, supports distributed memory, shared memory, multicore architectures, and uses climate and forecasts compliant Network Common Data Form for input/output. The package is freely available with an open source license from www.tendral.com/tsis/ .

Srinivasan, Ashwanth↗

Resolving the Evolution of Atomic Layer-Deposited Thin-Film Growth by Continuous In Situ X-Ray Absorption Spectroscopy

In situ synchrotron X-ray absorption near-edge structure characterization of thin-film titania growth by atomic layer deposition (ALD) over ZnO nanowires reveals persistent low-coordinated Ti motifs leading to a new picture of ALD growth. Through the design of growth and measurement cycles, Ti K-edge spectral data are continuously recorded so as to characterize the film evolution as a function of ALD cycle number and the surface changes within the time scale of the ALD cycle. A unified set of analysis tools is developed to interpret the time-series of spectral data. A prenucleation stage of growth, a transition region, and then a steady-state growth stage are observed with distinguishable features. Multivariate curve resolution analysis, that is physically constrained, demonstrates two specific spectral components with associated, time-dependent concentrations. The bulk-film component tracks the stages of growth. The surface and interface components, present throughout the stages of growth, reveal a significant coverage of relatively isolated or loosely networked tetrahedrally coordinated Ti atomic motifs. Lastly, spectral signatures for the intra-cycle growth kinetics are reconstructed at a time resolution of ~1 s and demonstrate that the transient Ti motifs on the growing surface stabilize within a few seconds of the Ti precursor pulse.

36 MATERIALS SCIENCE↗

Multivariate analysis: An essential for studying complex glasses

Understanding the impact of individual compositional components on the devitrification of complex multicomponent glasses, for example, 10–50+ oxides, typically requires numerous studies to examine each component's impact. Here we apply exploratory data analysis (EDA) to a heterogeneous data set of silicate glasses to determine the cations’ individual and interacting effects on the crystallization of nepheline (nominally NaAlSiO 4 ). Our data consisted of 795 simulated high-level nuclear waste glasses composed of, on average, 50 oxide components. We determine the interactions in the heterogeneous data that cause deviations from the behavior found in simplified composition studies. Using both univariate and bivariate EDA techniques, we demonstrate the importance of including calculated structural glass parameters on nepheline's devitrification, including field strength, cation-to-anion radius ratio, and single-bond strength. Here, we also show that studies with simplified glass compositions may fall short in generating knowledge directly transferrable to complex glass compositions. The method used in this study has the potential to inform experimental design for simplified compositions (~6+ oxides) that can generate knowledge directly transferrable to complex, multivariable compositions. The observations reported here have broad implications for any study attempting to map the physical properties of a complex glass containing numerous cations.

36 MATERIALS SCIENCE↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A New Approach to Evaluate and Reduce Uncertainty of Model-Based Biodiversity Projections for Conservation Policy Formulation

Biodiversity projections with uncertainty estimates under different climate, land-use, and policy scenarios are essential to setting and achieving international targets to mitigate biodiversity loss. Evaluating and improving biodiversity predictions to better inform policy decisions remains a central conservation goal and challenge. A comprehensive strategy to evaluate and reduce uncertainty of model outputs against observed measurements and multiple models would help to produce more robust biodiversity predictions. We propose an approach that integrates biodiversity models and emerging remote sensing and in-situ data streams to evaluate and reduce uncertainty with the goal of improving policy-relevant biodiversity predictions. In this work, we describe a multivariate approach to directly and indirectly evaluate and constrain model uncertainty, demonstrate a proof of concept of this approach, embed the concept within the broader context of model evaluation and scenario analysis for conservation policy, and highlight lessons from other modeling communities.

essential biodiversity variables↗

Time Series Classification for Locating Forced Oscillation Sources

Here, this article presents a machine learning based time-series classification method for using synchrophasor measurements to locate the source of forced oscillation (FO) for fast disturbance removal. First, multivariate time series (MTS) matrices are constructed by the most informative measurements selected by sequential feature selection from each power plant. Then, the Mahalanobis matrix is trained such that the Mahalanobis distance between the MTSs from the same class (i.e., with the same FO source location) are minimized and from different classes (i.e., with different FO source locations) are maximized. This allows MTSs to be classified by classifiers with class membership corresponding to the location of each FO source. To meet the runtime requirements of online matching, class templates are constructed to reduce data size and improve matching efficiency. To account for uncertainty in identifying the exact beginning of an FO event, dynamic time warping is used to align the out-of-sync MTSs. IEEE 39bus and WECC 179bus systems are used for algorithm development and validation. Simulation results demonstrate that the algorithm meets online operation runtime requirement with high accuracy using misaligned data sets.

42 ENGINEERING↗

High‐Resolution National‐Scale Water Modeling Is Enhanced by Multiscale Differentiable Physics‐Informed Machine Learning

Abstract The National Water Model (NWM) is a key tool for flood forecasting, planning, and water management. Key challenges facing the NWM include calibration and parameter regionalization when confronted with big data. We present two novel versions of high‐resolution (∼37 km 2 ) differentiable models (a type of hybrid model): one with implicit, unit‐hydrograph‐style routing and another with explicit Muskingum‐Cunge routing in the river network. The former predicts streamflow at basin outlets whereas the latter presents a discretized product that seamlessly covers rivers in the conterminous United States (CONUS). Both versions use neural networks to provide a multiscale parameterization and process‐based equations to provide a structural backbone, which were trained simultaneously (“end‐to‐end”) on 2,807 basins across the CONUS and evaluated on 4,997 basins. Both versions show great potential to elevate future NWM performance for extensively calibrated as well as ungauged sites: the median daily Nash‐Sutcliffe efficiency of all 4,997 basins is improved to around 0.68 from 0.48 of NWM3.0. As they resolve spatial heterogeneity, both versions greatly improved simulations in the western CONUS and also in the Prairie Pothole Region, a long‐standing modeling challenge. The Muskingum‐Cunge version further improved performance for basins >10,000 km 2 . Overall, our results show how neural‐network‐based parameterizations can improve NWM performance for providing operational flood predictions while maintaining interpretability and multivariate outputs. The modeling system supports the Basic Model Interface (BMI), which allows seamless integration with the next‐generation NWM. We also provide a CONUS‐scale hydrologic data set for further evaluation and use.

Song, Yalan [Civil and Environmental Engineering T↗

The Chemistry Graduate Student Experience: Findings from an ACS Survey

Graduate training is a key element in producing a scientific workforce that reflects the nation’s diversity. This paper examines data from a 2013 American Chemical Society (ACS) survey of 2,544 chemistry masters and doctoral students and reveals barriers to reaching this goal. Multivariate statistical analyses indicate that women reported significantly less supportive relationships with advisors. Women were less likely to plan to finish their degrees, and for PhD students, the discrepancy was larger for students at the start of their graduate program. Women were also less likely to pursue the next level of training, and the gender difference related to postdoctoral plans was greater for those who identified with a racial-ethnic group traditionally underrepresented in chemistry (underrepresented minority, URM). URM students who were beyond the first year of their graduate program reported significantly less supportive relationships with peers. They were also less likely to have funding sufficient to meet their needs and more often used personal resources including loans. Despite these difficulties, URM students were more likely to definitely plan to finish their degrees, and men who identified as URM were more likely to plan to pursue postdoctoral work. Independent of gender and identification as URMs, students in more highly ranked schools reported less advisor support. Extensive open-ended comments indicated that large proportions of the students desired more attention and meaningful feedback from advisors and changes within their programs to promote support for students and advisor accountability. Suggestions for future research are given, and a companion commentary discusses needed directions for change.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗