Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Multivariate discrimination in quantum target detection

In this work, we describe a simple multivariate technique of likelihood ratios for improved discrimination of signal and background in multi-dimensional quantum target detection. Furthermore, the technique combines two independent variables, time difference and summed energy, of a photon pair from the spontaneous parametric down-conversion source into an optimal discriminant. The discriminant performance was studied in experimental data and in Monte-Carlo modelling with clear improvement shown compared to previous techniques. As novel detectors become available, we expect this type of multivariate analysis to become increasingly important in multi-dimensional quantum optics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Search for $t\overline{t}$ resonances in fully hadronic final states in $pp$ collisions at $ \sqrt{s} $ = 13 TeV with the ATLAS detector

This paper presents a search for new heavy particles decaying into a pair of top quarks using 139 fb -1 of proton-proton collision data recorded at a centre-of-mass energy of √s = 13 TeV with the ATLAS detector at the Large Hadron Collider. The search is performed using events consistent with pair production of high-transverse-momentum top quarks and their subsequent decays into the fully hadronic final states. The analysis is optimized for resonances decaying into a $t\overline{t}$ pair with mass above 1.4 TeV, exploiting a dedicated multivariate technique with jet substructure to identify hadronically decaying top quarks using large-radius jets and evaluating the background expectation from data. No significant deviation from the background prediction is observed. Limits are set on the production cross-section times branching fraction for the new Z' boson in a topcolor-assisted-technicolor model. The Z' boson masses below 3.9 and 4.7 TeV are excluded at 95% confidence level for the decay widths of 1% and 3%, respectively.

Heavy quark production↗

AI-Enabled Operations at Fermi Complex: Multivariate Time Series Prediction for Outage Prediction and Diagnosis

The Main Control Room of the Fermilab accelerator complex continuously gathers extensive time-series data from thousands of sensors monitoring the beam. However, unplanned events such as trips or voltage fluctuations often result in beam outages, causing operational downtime. This downtime not only consumes operator effort in diagnosing and addressing the issue but also leads to unnecessary energy consumption by idle machines awaiting beam restoration. The current threshold-based alarm system is reactive and faces challenges including frequent false alarms and inconsistent outage-cause labeling. To address these limitations, we propose an AI-enabled framework that leverages predictive analytics and automated labeling. Using data from $2,703$ Linac devices and $80$ operator-labeled outages, we evaluate state-of-the-art deep learning architectures, including recurrent, attention-based, and linear models, for beam outage prediction. Additionally, we assess a Random Forest-based labeling system for providing consistent, confidence-scored outage annotations. Our findings highlight the strengths and weaknesses of these architectures for beam outage prediction and identify critical gaps that must be addressed to fully harness AI for transitioning downtime handling from reactive to predictive, ultimately reducing downtime and improving decision-making in accelerator management.

Jain, Milan [PNL, Richland] (ORCID:000000021676111↗

Monitoring Radiochemical Processing Streams for the 238 Pu Supply Program with Process Pulse II

Oak Ridge National Laboratory (ORNL) is developing advanced spectroscopic and real-time monitoring capabilities to improve the timeliness of analytical measurements and process decisions for the 238Pu Supply Program. Reducing the time, resources, and costs associated with each production campaign is critical because overlapping campaigns will be required to meet the production goals of the National Aeronautics and Space Administration. Real-time, in situ analytical measurements in the heavily shielded hot cells at the Radiochemical Engineering Development Center (REDC) will allow for rapid process information feedback and operational benefits that help the 238Pu supply program scale-up production efforts. Noteworthy steps were taken during Campaign 5 to establish the ability to monitor processing streams in real time with spectrophotometry and a commercially available online monitoring software called The Unscrambler X Process Pulse II (PP) multivariate statistical process monitoring system by Camo Analytics (version 5.60). PP automates univariate-type calculations within the software itself and executes multivariate models built using The Unscrambler X (version 10.4 or newer). The Unscrambler is a commercially available data analysis software made by the same company. PP is composed of easy-to-use-tools for all personnel, including data scientists and technicians. The software can be used to plot analyte concentration profiles, spectral data, and other process variables in real time. All process data are represented in a single view with interactive charts useful for viewing how a process evolves over time.

07 ISOTOPE AND RADIATION SOURCES↗

Determination of the uranium content of storage containers filled with waste from Fukushima using cosmic ray data

This report presents the results of an experimental study of cosmic ray tomography aimed at determining the uranium contents of waste containers filled with waste from Fukushima. Data, taken using the Giant Muon Tracker, were obtained on a set of scenes constructed from blocks of polyethylene, concrete, iron, and lead. The data have been provided to Toshiba and have been used to evaluate the composition of scenes using multivariate linear regression.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Identifying RR Lyrae Variable Stars in Six Years of the Dark Energy Survey

We present a search for RR Lyrae stars using the full six-year data set from the Dark Energy Survey covering ∼5000 deg2 of the southern sky. Using a multistage multivariate classification and light-curve template-fitting scheme, we identify RR Lyrae candidates with a median of 35 observations per candidate. We detect 6971 RR Lyrae candidates out to ∼335 kpc, and we estimate that our sample is >70% complete at ∼150 kpc. We find excellent agreement with other wide-area RR Lyrae catalogs and RR Lyrae studies targeting the Magellanic Clouds and other Milky Way satellite galaxies. We fit the smooth stellar halo density profile using a broken-power-law model with fixed halo flattening (q = 0.7), and we find strong evidence for a break at with an inner slope of and an outer slope of . We use our catalog to perform a search for Milky Way satellite galaxies with large sizes and low luminosities. Using a set of simulated satellite galaxies, we find that our RR Lyrae-based search is more sensitive than those using resolved stellar populations in the regime of large (r h ≳ 500 pc), low-surface-brightness dwarf galaxies. A blind search for large, diffuse satellites yields three candidate substructures. The first can be confidently associated with the dwarf galaxy Eridanus II. The second has a distance and proper motion similar to the ultrafaint dwarf galaxy Tucana II but is separated by ∼5 deg. The third is close in projection to the globular cluster NGC 1851 but is ∼10 kpc more distant and appears to differ in proper motion.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluating the potential of short-term instrument deployment to improve distributed wind resource assessment

Distributed wind projects, which are connected at the distribution level of an electricity system or in off-grid applications to serve specific or local energy needs, often rely solely on wind resource models to establish wind speed and energy generation expectations. Historically, anemometer loan programs have provided an affordable avenue for more accurate onsite wind resource assessment, and the lowering cost of lidar systems has shown similar advantages for more recent assessments. While a full 12 months of onsite wind measurement is the standard for correcting model-based long-term wind speed estimates for utility-scale wind farms, the time and capital investment involved in gathering onsite measurements must be reconciled with the energy needs and funding opportunities that drive expedient deployment of distributed wind projects. Much literature exists to quantify the performance of correcting long-term wind speed estimates with 1 or more years of observational data, but few studies explore the impacts of correcting with months-long observational periods. This study aims to answer the question of how short you can go in terms of the observational time period needed to make impactful improvements to model-based long-term wind speed estimates. Three algorithms, multivariable linear regression, adaptive regression splines, and regression trees, are evaluated for their skill at correcting long-term wind resource estimates from the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) using months-long periods of observational data from 66 locations across the US. On average, correction with even 1 month of observations provides significant improvement over the baseline ERA5 wind speed estimates and produces median bias magnitudes and relative errors within 0.22 m s −1 and 4 percentage points of the median bias magnitudes and relative errors achieved using the standard 12 months of data for correction. However, in cases when the shortest observational periods (1 to 2 months) used for correction are not well correlated with the overlapping ERA5 reference, the resultant long-term wind speed errors are worse than those produced using ERA5 without correction. Summer months, which are characterized by weaker relative wind speeds and standard deviations for most of the evaluation sites, tend to produce the worst results for long-term correction using months-long observations. The three tested algorithms perform similarly for long-term wind speed bias; however, regression trees perform notably worse than multivariable linear regression and adaptive regression splines in terms of correlation when using 6 months or less of observational data for correction. Translating the analysis to wind energy, median relative errors in the capacity factor are on average within 10 % using 1 month of training. If the observation period used for correction is not well correlated with the reference data, however, misrepresentation of the observed capacity factor can be substantial. The risk associated with poor correlation between the observed and reference datasets decreases with increasing training period length. In the worst-correlation scenarios, the median capacity factor relative errors from using 1, 3, and 6 months are within 47 %, 26 %, and 16 %, respectively.

17 WIND ENERGY↗

Crowd-based spatial risk assessment of urban flooding: Results from a municipal flood hotline in Detroit, MI

Climate change is increasing the frequency and intensity of extreme precipitation events, raising the risk of urban flood disasters. This study uses a crowd-sourced municipal call database to characterize the spatial distribution of flood risk in Detroit, MI. Call data including dates and addresses were obtained from the City of Detroit Department of Public Works for 2021. Calls were mapped and aggregated to census tract counts and merged with neighborhood-level data. Associations of predictors with flood calls were tested using spatial regression models. Flooding calls were located throughout the city but were concentrated in specific areas. Multivariate models of census tract level call counts indicated that increased poverty and Black, immigrant, and older residents were positively associated with flood calls, while increased elevation was associated with protective effects. Longer distances from waste water interceptors were associated with higher risk for calls. Crowd-sourced flood hotline call data can be used for effective spatial flood risk assessment. Though flooding occurs throughout the city of Detroit, infrastructural, neighborhood, and household factors influence flooding extent. Limitations included the self-reported nature of calls. Future modeling efforts might include input from local stakeholders to improve spatial risk assessment.

54 ENVIRONMENTAL SCIENCES↗

Impact of Precipitation Parameters on the Specific Surface Area of PuO 2

Controlling the properties of PuO 2 through processing is of vital importance to environmental transport and fate, production of nuclear fuels, nuclear forensic analyses, stockpile stewardship, and storage of nuclear wastes applications. A number of processing conditions have been identified to control final product properties, including specific surface area (SSA), residual carbon content, adsorption of volatile species, morphology, and particle size. In this paper, a novel approach is developed for the prediction of PuO 2 SSA via the synthetic route of Pu(IV) oxalate precipitation followed by calcination. The proposed model utilizes multivariate regression methodology and leave one out formalism to link Savannah River Site (SRS) precipitation and calcination production data to the SSA of the final product. A comparison among the models provides insight into the accuracy and ability to identify variations amongst the processing data. Additionally, the models may also be used to fit new data outside of the parameters explored in a production facility. Finally, the trained model was compared to a similarly trained conventional model form to illustrate the influence of precipitation parameters on the prediction of the final SSA. The models presented here attempt to provide new methods for more accurate prediction of the PuO 2 product properties in a production scale environment for key environmental and nuclear applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Data analytics for intermodal freight transportation applications

With the growth of intermodal freight transportation, it is important that transportation planners and decision-makers are knowledgeable about freight flow data to make informed decisions. This is particularly true with Intelligent Transportation Systems (ITS) offering new capabilities for intermodal freight transportation. Specifically, ITS enables access to multiple different data sources, but they have different formats, resolutions, and time scales. Thus, knowledge of data science is essential to be successful in future ITS-enabled intermodal freight transportation systems. This chapter discusses the commonly used descriptive and predictive data analytic techniques in intermodal freight transportation applications. These techniques cover the entire spectrum of univariate, bivariate, and multivariate analyses. In addition to illustrating how to apply these techniques manually, this chapter will also show how to apply them using the statistical software R. Additional exercises are provided for those who wish to apply the described techniques to more complex problems.

Huynh, Nathan↗

Modeling Stochastic Variability in Multiband Time-series Data

In preparation for the era of time-domain astronomy with upcoming large-scale surveys, we propose a state-space representation of a multivariate damped random walk process as a tool to analyze irregularly-spaced multifilter light curves with heteroscedastic measurement errors. We adopt a computationally efficient and scalable Kalman filtering approach to evaluate the likelihood function, leading to maximum O(k 3 n) complexity, where k is the number of available bands and n is the number of unique observation times across the k bands. This is a significant computational advantage over a commonly used univariate Gaussian process that can stack up all multiband light curves in one vector with maximum O(k 3 n 3 ) complexity. Using such efficient likelihood computation, we provide both maximum likelihood estimates and Bayesian posterior samples of the model parameters. Three numerical illustrations are presented: (i) analyzing simulated five-band light curves for a comparison with independent single-band fits; (ii) analyzing five-band light curves of a quasar obtained from the Sloan Digital Sky Survey Stripe 82 to estimate short-term variability and timescale; (iii) analyzing gravitationally lensed g- and r-band light curves of Q0957+561 to infer the time delay. Two R packages, Rdrw and timedelay, are publicly available to fit the proposed models.

79 ASTRONOMY AND ASTROPHYSICS↗

Landmark-Warped Emulators for Models with Misaligned Functional Response

Many computer models output functional data, and in some cases, these functional data have similar, but misaligned, shape characteristics. In this paper, we introduce a general approach for building emulators for computer models that output misaligned functional data when key values in the functional response (landmarks) can be easily identified. This approach has two main parts: modeling the aligned (using the landmarks) functional data, and modeling the functions that map the misaligned data to the aligned space (warping functions). As the warping functions are required to be monotonic, we give special attention to modeling monotonic functional response data. We discuss how our approach can be easily applied for a variety of typical emulators, such as Gaussian processes, Bayesian multivariate adaptive regression splines, and Bayesian additive regression trees, and how sensitivity analysis can be performed. We demonstrate our approach by building emulators for two applications: (1) a high-energy-density physics computer model used to simulate inertial confinement fusion ignition experiments, where model outputs are highly misaligned, and (2) a multiphysics continuum hydrocode used to simulate high-velocity impact experiments, where model outputs are only slightly misaligned. In case (1) traditional methods cannot be applied, while in (2) they can be applied, but the proposed method performs significantly better.

97 MATHEMATICS AND COMPUTING↗

Data to Accompany: PM2.5 is insufficient to explain personal PAH exposure

Fine particulate matter (PM2.5) air quality index (AQI) data from outdoor stationary monitors and Hazard Mapping System (HMS) smoke density data from satellites are often used as proxies for personal chemical exposure. Silicone wristbands can quantify more individualized exposure data than stationary air monitors or smoke satellites. However, it is not understood how these proxy measurements compare to chemical data measured from wristbands. We hypothesized that predictive models for personal chemical exposure would be significantly improved by expanding beyond stationary PM2.5 AQI data or satellite HMS data to also include environmental and behavioral information. In Eugene, Oregon, participants wore daily wristbands, carried a phone that recorded locations, and answered daily questionnaires for a seven-day period in multiple seasons. We gathered publicly available daily PM2.5 AQI data and HMS data. We analyzed wristbands for 94 organic chemicals, including 53 polycyclic aromatic hydrocarbons (PAHs). Wristband chemical detections and concentrations, behavioral variables (e.g., time spent indoors), and environmental conditions (e.g., PM2.5 AQI) significantly differed between seasons. Machine learning models were fit to predict personal chemical exposure using PM2.5 AQI only, HMS only, and a multivariate feature set including PM2.5 AQI, HMS, and other environmental and behavioral information. On average, the multivariate models increased predictive accuracy by approximately 70% compared to either the AQI model or the HMS model for all chemicals modeled. This study provides evidence that PM2.5 AQI data alone or HMS data alone is insufficient to explain personal chemical exposures. Our results identify additional key predictors of personal chemical exposure.

Bramer, Lisa M↗

Deficient precipitation sensitivity to Sahel land surface forcings among CMIP5 models

Abstract The overall performance of the simulated seasonal precipitation response to local terrestrial forcings, namely vegetation abundance and soil moisture, in the Sahel among the Coupled Model Intercomparison Project Phase Five (CMIP5) Earth System Models (ESMs) is systematically investigated and compared with its observational counterpart using a multivariate statistical method. The observed seasonal precipitation response is evaluated against a large ensemble of observational, reanalysis, and satellite data sets to provide quantification of uncertainties. The behaviour of models with and without a Dynamic Global Vegetation Model (DGVM) component is also explored, along with the mechanisms responsible for terrestrial feedback on rainfall. In general, the CMIP5 models can reasonably capture the seasonal evolution of Sahel precipitation and soil moisture, albeit with wet biases during the pre‐monsoon period and dry biases during the peak monsoon period. The non‐DGVM ESMs simulate comparable leaf area indices (LAIs) with observations, while DGVM‐enabled ESMs simulate too much year‐round LAI. The variance of precipitation that is attributed to oceanic forcings in CMIP5 is comparable with observations; however, the variance of precipitation that is attributed to terrestrial forcings is smaller in CMIP5 models than observed, especially for non‐DGVM ESMs. CMIP5 models, especially those without DGVMs, undervalue precipitation's observed response strength to soil moisture anomalies. In both observations and CMIP5 models, none of the atmospheric variables show significant responses to direct vegetation forcing, except for the response in transpiration. Although vegetation has minimal direct effect on the atmospheric state, it can affect the atmosphere by modifying soil moisture and transpiration rate indirectly, which helps explain the more realistic simulation of rainfall in DGVM‐enabled ESMs than non‐DGVM ESMs. Coupling of an ESM to a DGVM is critical in generating reasonable land–atmosphere feedback and examining future ecological and climatic changes over the Sahel.

54 ENVIRONMENTAL SCIENCES↗

A Deep Neural Network for Simultaneous Estimation of b Jet Energy and Resolution

We describe a method to obtain point and dispersion estimates for the energies of jets arising from b quarks produced in proton–proton collisions at an energy of $\sqrt{s}=13\,\text {TeV} $ at the CERN LHC. The algorithm is trained on a large sample of simulated b jets and validated on data recorded by the CMS detector in 2017 corresponding to an integrated luminosity of 41 $\,\text {fb}^{-1}$. A multivariate regression algorithm based on a deep feed-forward neural network employs jet composition and shape information, and the properties of reconstructed secondary vertices associated with the jet. The results of the algorithm are used to improve the sensitivity of analyses that make use of b jets in the final state, such as the observation of Higgs boson decay to $\hbox {b}\bar{\hbox {b}}$.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Advanced Visualization for Scientific Data Analysis and Insight [Slides]

This talk will explore how we have used advanced visualization technologies to support analytical reasoning and knowledge discovery. Specifically, we will present several examples detailing some recent scientific successes using state-of-the-art immersive and high-resolution visualization at the National Renewable Energy Laboratory's Computational Science Center. On multiple occasions, we have observed scientists and engineers discover features in their data using advanced visualization technologies that they had not seen in prior investigations of their data on traditional desktop displays. We have embedded more information into our analytics tools, allowing engineers to explore complex multivariate spaces. We have observed how interactions seem to catalyze understanding.

97 MATHEMATICS AND COMPUTING↗

Detection of Synchrophasor False Data Injection Attack using Feature Interactive Network

The synchrophasor data recorded by Phasor Measurement Units (PMUs) plays an increasingly critical role in the regulation and situational awareness of power systems. However, the widely installed PMUs are vulnerable to multiple malicious attacks from cyber hackers during data transmission and storage. To address this problem, a Modified Ensemble Empirical Mode Decomposition (MEEMD) is proposed first to extract the intrinsic mode functions of each Synchrophasor Data Attacks (SDA). The frequency-based adaptive screening criterion embedded in MEEMD is used to eliminate the false intrinsic mode functions. Next, a Multivariate Convolutional Neural Network (MCNN) is proposed to identify multiple SDA by utilizing the extracted intrinsic mode functions and original SDA as input vectors. A fusion block as the main structure of MCNN is also leveraged to increase the diversity of features and compress the model parameters. Integrating MEEMD and MCNN, a framework with automatic feature extraction and multi-source information fusion capability, referred to as Feature Interactive Network (FIN), is proposed to detect multiple SDA. Based on the proposed FIN framework, six types of SDA are explored for the first time using actual synchrophasor data in FNET/Grideye that was collected from different locations in the U.S. Eastern Interconnection. Finally, a large quantity of experiments with different attack strengths are used to evaluate the adaptability and classification performance of the proposed FIN.

24 POWER TRANSMISSION AND DISTRIBUTION↗