Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

CTGAN-TVAE

SAND2026-18914O CTGAN-TVAE (Conditional Tabular Generative Adversarial Networks-Tabular Variational Autoencoders) generates extensive sets of variable generation data through a hybrid framework. It enhances latent space representation by combining TVAE's robust feature-embedding with CTGAN's ability to condition categorical variables such as time. CTGAN-TVAE employs a fully connected neural network within a conditional generative adversarial network framework to manage continuous and categorical data effectively, capturing complex feature interactions without needing sequential modeling. This was developed as part of NNSA-MSIPP: Minority Serving Institution Partnership Program, Grant Number DE-NA0004016. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Newlun, Cody [Sandia National Lab. (SNL-CA), Liver↗

2000-2001 California Statewide Household Travel Survey

The 2000-2001 California Statewide Household Travel Survey, conducted under the auspices of the California Department of Transportation, was conducted between October 2000 and December 2001 among households located in each of California's 58 counties. The purpose of the study was to update the statewide database of household socioeconomic and travel information and helped to refine travel estimates, models, and forecasts throughout California. The survey was an essential element in determining statewide and regional travel patterns. A total of 17,040 households participated in the survey, contributing household socioeconomic and travel data. Travel variables collected include trip times, mode, activity at location, origin and destination, and vehicle occupancy, among other travel-related data from 134,173 trips. Residents completed diary records of their daily travel over a 24-hour (weekday) or 48-hour period (Friday/Saturday or Sunday/Monday pair).

1Hz data↗

Robust errant beam prognostics with conditional modeling for particle accelerators

Abstract Particle accelerators are complex and comprise thousands of components, with many pieces of equipment running at their peak power. Consequently, they can fault and abort operations for numerous reasons, lowering efficiency and science output. To avoid these faults, we apply anomaly detection techniques to predict unusual behavior and perform preemptive actions to improve the total availability. Supervised machine learning (ML) techniques such as siamese neural network models can outperform the often-used unsupervised or semi-supervised approaches for anomaly detection by leveraging the label information. One of the challenges specific to anomaly detection for particle accelerators is the data’s variability due to accelerator configuration changes within a production run of several months. ML models fail at providing accurate predictions when data changes due to changes in the configuration. To address this challenge, we include the configuration settings into our models and training to improve the results. Beam configurations are used as a conditional input for the model to learn any cross-correlation between the data from different conditions and retain its performance. We employ conditional siamese neural network (CSNN) models and conditional variational auto encoder (CVAE) models to predict errant beam pulses at the spallation neutron source under different system configurations and compare their performance. We demonstrate that CSNNs outperform CVAEs in our application.

43 PARTICLE ACCELERATORS↗

A Graphical Model for Fusing Diverse Microbiome Data

This paper develops a Bayesian graphical model for fusing disparate types of count data. The motivating application is the study of bacterial communities from diverse high-dimensional features, in this case, transcripts, collected from different treatments. In such datasets, there are no explicit correspondences between the communities and each corresponds to different factors, making data fusion challenging. We introduce a flexible multinomial-Gaussian generative model for jointly modeling such count data. This latent variable model jointly characterizes the observed data through a common multivariate Gaussian latent space that parameterizes the set of multinomial probabilities of the transcriptome counts. The covariance matrix of the latent variables induces a covariance matrix of co-dependencies between all the transcripts, effectively fusing multiple data sources. We present a computationally scalable variational Expectation-Maximization (EM) algorithm for inferring the latent variables and the parameters of the model. Here, the inferred latent variables provide a common dimensionality reduction for visualizing the data and the inferred parameters provide a predictive posterior distribution. In addition to simulation studies that demonstrate the variational EM procedure, we apply our model to a bacterial microbiome dataset.

59 BASIC BIOLOGICAL SCIENCES↗

Experimental Analysis of the Effects of Simulator Complexity on Human Performance

Human Reliability Analysis predicts accidents caused by human errors and is an important factor in Probabilistic Safety Assessment that comprehensively evaluates the safety of nuclear power plants. This study compares and analyzes the human performance of nuclear power plant operators according to the simulator complexity as part of the HRA data collection support method development project conducted by Idaho National Laboratory. This experiment was conducted by setting two types of simulators and scenarios as independent variables. The data collected by conducting the experiments in two different simulators were analyzed using an analysis of variance test and a correlation analysis, and four human performance charts were derived.

99 GENERAL AND MISCELLANEOUS↗

Opportunities for Using the Industrial Assessment Center Database for Industrial Water Use Analysis

The manufacturing sector accounted for approximately 5–6% of total U.S. water use in 2015. Of that amount, 75–80% is self supplied withdrawal from surface-water and groundwater sources and the remainder is from public water supplies. Although manufacturing facilities commonly locate in water-scarce areas, water scarcity still poses a great risk to the manufacturing sector. Reliable water is necessary for any facility that relies on it for process and comfort cooling, cleaning, employee use, and steam generation. One of these barriers to water efficiency is the lack of reliable data on overall U.S. industrial water use—how it is used and the quantities required for each sector. If a facility cannot be easily compared with a facility of similar size and sector, knowing if it is effectively using water conservation best practices is difficult. One potential source of industrial water use data is the U.S. Department of Energy (DOE)–sponsored Industrial Assessment Centers (IACs). IACs are university-based organizations that provide free audits to small- and medium-sized manufacturing facilities to identify productivity improvement and waste and energy reduction opportunities. The IACs also maintain a database of all the audits conducted, which currently holds more than 19,267 assessments and 145,000 recommendations (as of July 24, 2020). This database also contains energy utility (electricity, natural gas, and other fuels) and water utility data, making it a potential data source for industrial water use. This report attempts to create regression models to predict a small- or medium-sized industrial facility’s annual water use or cost based on its industrial subsector and several possible relevant variables. Using data collected by IAC assessments, models for several industrial subsectors were generated via stepwise regression techniques to determine which variables (annual sales, number of employees, facility/plant area, annual production hours, and a water stress metric) are relevant.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Anomalous structural recovery in the near glass transition range in a polymer glass: Data revisited in light of temperature variability in vacuum oven‐based experiments*

Abstract There is considerable interest in data reported by Cangialosi et al. [Cangialosi, D., Boucher, V. M., Alegría, A., & Colmenero, J. (2013). Physical Review Letters , 111 (9), 095701] that purportedly showed unusual irregular responses in the isothermal structural recovery behavior of glasses because the observation challenges the general view of the dynamics being smooth functions of time as observed in extensive dilatometric experiments by Kovacs [Kovacs, A. J. (1964). In Fortschritte der Hochpolymeren‐Forschung , 3 , 394–507. Springer, Berlin, Heidelberg.] in the 1960s. The reported data show an apparent two mechanism structural (enthalpy) recovery at aging temperatures ranging from 6 to 17 K below which is also controversial as it could not be reproduced in work from Koh and Simon [Koh, Y. P., & Simon, S. L. (2013). Macromolecules , 46(14), 5815–5821.] where a second plateau observed in the enthalpy loss curve was not reproduced in nearly 1‐year‐long experiments at 15 K below . Therefore, it is important to determine what might be possible reasons for the apparent non‐smooth relaxations. Here, we examine the possibility that the poor temperature control of a typical vacuum oven could explain the anomalous results. We used the Tool–Narayanaswamy–Moynihan (TNM) model of structural recovery to calculate the effect of typical vacuum oven temperature variation on the structural recovery of polystyrene and find that we can reproduce the reported experimental results. The issue of temperature control using a vacuum oven is discussed, and data from thermocouple measurements of temperature at various locations inside a vacuum oven are shown.

36 MATERIALS SCIENCE↗

Data for Greenhouse Gas Accounting Procedures in Low Carbon Fuel Policies Overlook the Spatial Variability of Miscanthus-Derived Sustainable Aviation Fuel

Low carbon fuel policies such as the U.S. Renewable Fuel Standard (RFS), Canada Clean Fuel Regulations (CFR), and California Low Carbon Fuel Standard (LCFS) as well as the 45Z tax credit are intended to reduce greenhouse gas (GHG) emissions from transportation. Cellulosic feedstocks, optimized biorefineries, and favorable farming locations can significantly reduce biofuel carbon intensity (CI). Despite advances in field-to-fuel GHG monitoring and flexibility in resource allocation within biorefineries (e.g., governing net electricity production), rigid CI accounting procedures in current policies may limit CI responsiveness across candidate sites and processing facilities. This work examines a hypothetical biomass-to-sustainable aviation fuel (SAF) pathway using miscanthus and alcohol-to-jet (i) to demonstrate how GHG accounting requirements drive estimates of biofuel CIs and (ii) to explore potential CI and financial implications of scenario-specific life cycle assessment (LCA). Results demonstrate that GHG accounting using the CFR/LCFS can reasonably account for distinct levels of net electricity production by a biorefinery, but only the CFR yields similar CI sensitivity to spatially explicit factors (feedstock CI, grid electricity CI) as scenario-specific LCA: most GHG accounting frameworks do not capture CI variation across candidate sites in the United States. Ultimately, this work demonstrates the importance of LCA methodological specifications in low carbon fuel policies and tax credits.

Miscanthus↗

Replication Data for: Multitask Machine Learning of Collective Variables for Enhanced Sampling of Rare Events

The data underlying this published work have been made publicly available in this repository as part of the IMASC Data Management Plan. This work was supported as part of the Integrated Mesoscale Architectures for Sustainable Catalysis (IMASC), an Energy Frontier Research Center funded by the U.S. Department of Energy, Office of Science, Basic Energy Sciences under Award # DE-SC0012573.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Aboveground biomass density models for NASA’s Global Ecosystem Dynamics Investigation (GEDI) lidar mission

NASA's Global Ecosystem Dynamics Investigation (GEDI) is collecting spaceborne full waveform lidar data with a primary science goal of producing accurate estimates of forest aboveground biomass density (AGBD). This paper presents the development of the models used to create GEDI's footprint-level (~25 m) AGBD (GEDI04_A) product, including a description of the datasets used and the procedure for final model selection. The data used to fit our models are from a compilation of globally distributed spatially and temporally coincident field and airborne lidar datasets, whereby we simulated GEDI-like waveforms from airborne lidar to build a calibration database. We used this database to expand the geographic extent of past waveform lidar studies, and divided the globe into four broad strata by Plant Functional Type (PFT) and six geographic regions. GEDI's waveform-to-biomass models take the form of parametric Ordinary Least Squares (OLS) models with simulated Relative Height (RH) metrics as predictor variables. From an exhaustive set of candidate models, we selected the best input predictor variables, and data transformations for each geographic stratum in the GEDI domain to produce a set of comprehensive predictive footprint-level models. We found that model selection frequently favored combinations of RH metrics at the 98th, 90th, 50th, and 10th height above ground-level percentiles (RH98, RH90, RH50, and RH10, respectively), but that inclusion of lower RH metrics (e.g. RH10) did not markedly improve model performance. Second, forced inclusion of RH98 in all models was important and did not degrade model performance, and the best performing models were parsimonious, typically having only 1-3 predictors. Third, stratification by geographic domain (PFT, geographic region) improved model performance in comparison to global models without stratification. Fourth, for the vast majority of strata, the best performing models were fit using square root transformation of field AGBD and/or height metrics. There was considerable variability in model performance across geographic strata, and areas with sparse training data and/or high AGBD values had the poorest performance. These models are used to produce global predictions of AGBD, but will be improved in the future as more and better training data become available.

54 ENVIRONMENTAL SCIENCES↗

A Framework for Closed-Loop Optimization of an Automated Mechanical Serial-Sectioning System via Run-to-Run Control as Applied to a Robo-Met.3D

Optimization of automated data collection is gaining increased interest for the purposes of enabling closed-loop self-correcting systems that inherently maximize operational efficiencies and reduce waste. Many data collection systems have several variables which influence data accuracy or consistency and which can require frequent user interaction to be monitored and maintained. Operating upon a Robo-MET.3D™ automated mechanical serial-sectioning system, a run-to-run control algorithm has been developed to accelerate data collection and reduce data inconsistency. Here, using historical data amassed over a decade of experiments, a linear regression model of the deterministic system dynamics is created and used to employ a run-to-run control algorithm that optimizes selected system inputs to reduce operator intervention and increase efficacy while reducing variance of system output.

42 ENGINEERING↗

Data-driven analysis of relight variability of jet fuels induced by turbulence

For safety purposes, reliable reignition of aircraft engines in the event of flame blow-out is a critical requirement. Typically, an external ignition source in the form of a spark is used to achieve a stable flame in the combustor. However, such forced turbulent ignition may not always successfully relight the combustor, mainly because the state of the combustor cannot be precisely determined. Uncertainty in the turbulent flow inside the combustor, inflow conditions, and spark discharge characteristics can lead to variability in sparking outcomes even for nominally identical operating conditions. Prior studies have shown that of all the uncertain parameters, turbulence is often dominant and can drastically alter ignition behavior. For instance, even when different fuels have similar ignition delay times, their ignition behavior in practical systems can be completely different. In practical operating conditions, it is challenging to understand why ignition fails and how much variation in outcomes can be expected. The focus of this work is to understand relight variability induced by turbulence for two different aircraft fuels, namely Jet-A and a variant named C1. A detailed, previously developed simulation approach is used to generate a large number of successful and failed ignition events. Using this data, the cause of misfire is evaluated based on a discriminant analysis that delineates the difference between turbulent initial conditions that lead to ignition or failure. From the discriminant analysis, a compressed sensing algorithm is then applied to help pinpoint the locations of relevant turbulent features. Findings from the discriminant analysis are confirmed with the time history of near kernel properties. Next, a clustering strategy is used to identify ignition and misfire modes. With this approach, it was determined that the cause of ignition failure is different for the two fuels. While it was found that Jet-A is influenced by fuel entrainment, C1 was found to be more sensitive to small scale turbulence features. Finally, a larger variability is found in the ignition modes of C1, which can be subject to extreme events induced by kernel breakdown.

42 ENGINEERING↗

Floods and Heavy Precipitation at the Global Scale: 100‐Year Analysis and 180‐Year Reconstruction

Floods and heavy precipitation have disruptive impacts worldwide, but their historical variability remains only partially understood at the global scale. This article aims at reducing this knowledge gap by jointly analyzing seasonal maxima of streamflow and precipitation at more than 3,000 stations over a 100‐year period. The analysis is based on Hidden Climate Indices (HCIs). Like standard climate indices (e.g., Nino 3.4, NAO), HCIs are used as covariates explaining the temporal variability of data, but unlike them, HCIs are estimated from the data. In this work, a distinction is made between common HCIs, that affect both heavy precipitation and floods, and specific HCIs, that exclusively affect one or the other. Overall, HCIs do not show noticeable autocorrelation, but some are affected by noticeable trends. In particular, strong and wide‐ranging trends are identified in precipitation‐specific HCIs, while trends affecting flood‐specific HCIs are weaker and have more localized effects. A probabilistic model is then derived to link HCIs and large‐scale atmospheric variables (pressure, wind, temperature) and to reconstruct HCIs since 1836 using the 20CRv3 reanalysis. In turn this allows estimating the probability of occurrence of floods and heavy precipitation at the global scale. This 180‐year reconstruction highlights flood hot‐spots and hot‐moments in the distant past, well before the establishment of perennial monitoring networks. Finally, the approach presented in this study is generic and paves the way for an improved characterization of historical variability by making a better use of long but highly irregular station data sets.

54 ENVIRONMENTAL SCIENCES↗

NeuroSEM: A hybrid framework for simulating multiphysics problems by coupling PINNs and spectral elements

Multiphysics problems that are characterized by complex interactions among fluid dynamics, heat transfer, structural mechanics, and electromagnetics, are inherently challenging due to their coupled nature. While experimental data on certain state variables may be available, integrating these data with numerical solvers remains a significant challenge. Physics-informed neural networks (PINNs) have shown promising results in various engineering disciplines, particularly in handling noisy data and solving inverse problems in partial differential equations (PDEs). However, their effectiveness in forecasting nonlinear phenomena in multiphysics regimes, particularly involving turbulence, is yet to be fully established. Here, this study introduces NeuroSEM, a hybrid framework integrating PINNs with the highfidelity Spectral Element Method (SEM) solver, Nektar++. NeuroSEM leverages the strengths of both PINNs and SEM, providing robust solutions for multiphysics problems. PINNs are trained to assimilate data and model physical phenomena in specific subdomains, which are then integrated into the Nektar++ solver. We demonstrate the efficiency and accuracy of NeuroSEM for thermal convection in cavity flow and flow past a cylinder. The framework effectively handles data assimilation by addressing those subdomains and state variables where the data is available. We applied NeuroSEM to the Rayleigh-B´enard convection system, including cases with missing thermal boundary conditions and noisy datasets. Finally, we applied the proposed NeuroSEM framework to real particle image velocimetry (PIV) data to capture flow patterns characterized by horseshoe vortical structures. Our results indicate that NeuroSEM accurately models the physical phenomena and assimilates the data within the specified subdomains. The framework’s plug-and-play nature facilitates its extension to other multiphysics or multiscale problems. Furthermore, NeuroSEM is optimized for efficient execution on emerging integrated GPU-CPU architectures. This hybrid approach enhances the accuracy and efficiency of simulations, making it a powerful tool for tackling complex engineering challenges in various scientific domains.

42 ENGINEERING↗

Total Dissolved Nitrogen and Ammonia Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for total dissolved nitrogen (TDN) and ammonia concentrations for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. TDN was analyzed using a Shimadzu Total Nitrogen Module (TNM-1) combined with the TOC-VCSH analyzer (Shimadzu Corporation, Japan). TNM-1 is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. Ammonia was determined using a Lachat's QuikChem 8500 Series 2 Flow Injection Analysis System (LACHAT Instruments, QuckChem 8500 series 2, Automated Ion Analyzer, Loveland, Colorado). When ammonia in water samples is heated (60 degrees C) with salicylate and hypochlorite in an alkaline phosphate buffer, an emerald green color is produced which is proportional to the ammonia concentration. The color is intensified by the addition of nitroprusside. Ethylenediaminetetraacetic acid (EDTA) is added to the buffer to prevent the interference of metal ions (Ca, Mg, and Fe etc.). Ammonia-N is then determined by LACHAT flow injection and a colorimetric assay at an absorbance wavelength 660 nm. (Reference: LACHAT Instruments: QuickChem Method 90-107-06-3-A, Determination of Ammonia by Flow Injection Analysis (High Throughput, Salicylate Method/DCIC) (Multi Matrix method). Written by Lynn Egan (Application group), February 08, 2011.) All files are labeled by location and variable, and data reported are the mean values upon replicate measurements. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process as detailed in the methods. This data package contains (1) a zip file (tdn_ammonia_data_2015-2025.zip) containing a total of 299 files: 298 data files of ammonia and TDN data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the determination of Method Detection Limits (MDLs) for TDN data, which has been updated in 2026-08; and (5) PDF and docx files for the detemination of Method Detection Limits (MDLs) for Ammonia and the Interferences by LACHAT Flow Injection Analysis. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 105 locations containing TDN and Ammonia-N data. Update 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses and Determination of Method Detection Limit for Ammonia and the Interferences by LACHAT Flow Injection Analysis documents, which can be accessed as PDFs or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated total dissolved nitrogen and ammonia data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Units were listed incorrectly, but have been fixed to reflect correct units (ug/L). File level metadata (flmd) and data dictionary (dd) files were updated to reflect the updated versions of these files. Available data was added up until 2022-06-01. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-27. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for TDN were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for TDN data were added to this dataset.

54 ENVIRONMENTAL SCIENCES↗

Review of Technical Photovoltaic Key Performance Indicators and the Importance of Data Quality Routines

Technical key performance indicators (KPIs) are important metrics used to assess and quantitatively summarize various aspects of photovoltaic (PV) systems, including long-term performance, economic viability, and carbon footprint. Herein, a group of experts of the International Energy Agency's Photovoltaic Power Systems Programme Task 13 collect and describ the most important technical KPIs used in the industry. Thereby, a set of best practices for reliably handling PV system data is presented and the impact of data quality and climatic variability on KPI calculation is investigated. Further, the effective use of technical KPIs allows triggering data-driven and informed decisions to optimize PV systems and providing a comprehensive overview of how PV systems operate across different conditions and climates. With the worldwide growth of the PV industry, more companies operate/own PV systems in different regions, where the climatic and seasonal profiles differ. This requires context-aware evaluation of KPIs, or the judicious application of multiple KPIs, to ensure that each asset is evaluated correctly. Beyond that, there is untapped potential in the utilization of KPIs through geospatial mapping and extrapolation of fleet KPIs. This study demonstrates that the uncertainty in KPI estimation is not well understood and depends on data quality, climatic variability, and system configuration.

14 SOLAR ENERGY↗

Joint Modeling of Quasar Variability and Accretion Disk Reprocessing Using Latent Stochastic Differential Equations

Quasars are bright active galactic nuclei powered by the accretion of matter around supermassive black holes at the center of galaxies. Their stochastic brightness variability depends on the physical properties of the accretion disk and black hole. The upcoming Rubin Observatory Legacy Survey of Space and Time (LSST) is expected to observe tens of millions of quasars, so there is a need for efficient techniques like machine learning that can handle the large volume of data. Quasar variability is believed to be driven by an X-ray corona, which is reprocessed by the accretion disk and emitted as UV/optical variability. We are the first to introduce an auto-differentiable simulation of the accretion disk and reprocessing. We use the simulation as a direct component of our neural network to jointly model the driving variability and reprocessing, trained with supervised learning on simulated LSST-like 10 yr quasar light curves. We encode the light curves using a transformer encoder, and the driving variability is reconstructed using latent stochastic differential equations, a physically motivated generative deep learning method that can model continuous-time stochastic dynamics. By embedding the physical processes of the driving signal and reprocessing into our network, we achieve a model that is more robust and interpretable. We demonstrate that our model outperforms a Gaussian process regression baseline and can infer accretion disk parameters and time delays between wave bands, even for out-of-distribution driving signals. Our approach provides a powerful framework that can be adapted to solve other inverse problems in multivariate time series.

Fagin, Joshua [City Univ. of New York (CUNY), NY (↗

Multi-fidelity Bayesian neural networks: Algorithms and applications

Here we propose a new class of Bayesian neural networks (BNNs) that can be trained using noisy data of variable fidelity, and we apply them to learn function approximations as well as to solve inverse problems based on partial differential equations (PDEs). These multi-fidelity BNNs consist of three neural networks: The first is a fully connected neural network, which is trained following the maximum a posteriori probability (MAP) method to fit the low-fidelity data; the second is a Bayesian neural network employed to capture the cross-correlation with uncertainty quantification between the low- and high-fidelity data; and the last one is the physics-informed neural network, which encodes the physical laws described by PDEs. For the training of the last two neural networks, we first employ the mean-field variational inference (VI) to maximize the evidence lower bound (ELBO) to obtain informative prior distributions for the hyperparameters in the BNNs, and subsequently we use the Hamiltonian Monte Carlo (HMC) method to estimate accurately the posterior distributions for the corresponding hyperparameters. We demonstrate the accuracy of the present method using synthetic data as well as real measurements. Specifically, we first approximate a one- and four-dimensional function, and then infer the reaction rates in one- and two-dimensional diffusion-reaction systems. Moreover, we infer the sea surface temperature (SST) in the Massachusetts and Cape Cod Bays using satellite images and in-situ measurements. Taken together, our results demonstrate that the present method can capture both linear and nonlinear correlation between the low- and high-fidelity data adaptively, identify unknown parameters in PDEs, and quantify uncertainties in predictions, given a few scattered noisy high-fidelity data. Finally, we demonstrate that we can effectively and efficiently reduce the uncertainties and hence enhance the prediction accuracy with an active learning approach, using as examples a specific one-dimensional function approximation and an inverse PDE problem.

97 MATHEMATICS AND COMPUTING↗