Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gradient information”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Residential Demand Side Aggregation of Privacy-Conscious Consumers

The increasing adoption of smart meters has led to growing concerns regarding privacy risks stemming from the high resolution measurements. This has given rise to privacy protection techniques that physically alter the consumer's energy load profile, masking private information by using localised devices, e.g. batteries or flexible loads. Meanwhile, there has also been increasing interest in aggregating the distributed energy resources (DERs) of residential consumers to provide services to the grid. In this paper, we propose an online distributed algorithm to aggregate the DERs of privacy-conscious consumers to provide services to the grid, whilst preserving their privacy. Results show that the optimisation solution from the distributed method converges to one close to the optimum computed using an ideal centralised solution method, balancing between grid service provision, consumer preferences and privacy protection. More importantly, the distributed method preserves consumer privacy, and does not require high-bandwidth two-way communications infrastructure.

ancillary services↗

Comparison of EBSD, DIC, AFM, and ECCI for active slip system identification in deformed Ti-7Al

The use of SEM-DIC, AFM, ECCI, and HR-EBSD to characterize slip-system activity was assessed on the same material volume of Ti-7Al. This study presents a robust comparison of the various methods for the first time, including an assessment of their advantages and disadvantages, and how they can be used effectively in a complementary fashion. The analysis of the different approaches was carried out in a blind, round-robin manner at three different universities. A Ti-7Al specimen was deformed in uniaxial tension to approximately 3% axial strain, and the active slip systems were independently identified using (i) trace analysis; (ii) in-SEM digital image correlation, (iii) observations of residual dislocations from ECCI, and (iv) long-range rotation gradients through HR-EBSD, with consistent trace identification in all cases. Displacement data from AFM was used to augment SEM-DIC displacement data by providing complementary out-of-plane displacement information. Furthermore, short-range dislocation gradients (measured by DIC) provided insight into the residual geometrically necessary dislocation (GND) content, and was consistent with the GND content extracted from EBSD data and ECCI images, confirming the presence of residual GNDs on the dominant slip systems resulting in visible slip bands. Furthermore, these approaches can be used in tandem to provide multi-modal information on slip band identification, strain and orientation gradients, out-of-plane displacements, and the presence of GNDs and SSDs, all of which can be used to inform and validate the development of dislocation-based crystal plasticity and strain gradient models.

36 MATERIALS SCIENCE↗

SoDaH: the SOils DAta Harmonization database, an open-source synthesis of soil data from research networks, version 1.0

Data collected from research networks present opportunities to test theories and develop models about factors responsible for the long-term persistence and vulnerability of soil organic matter (SOM). Synthesizing datasets collected by different research networks presents opportunities to expand the ecological gradients and scientific breadth of information available for inquiry. Synthesizing these data is challenging, especially considering the legacy of soil data that have already been collected and an expansion of new network science initiatives. To facilitate this effort, here we present the SOils DAta Harmonization database (SoDaH; https://lter.github.io/som-website, last access: 22 December 2020), a flexible database designed to harmonize diverse SOM datasets from multiple research networks. SoDaH is built on several network science efforts in the United States, but the tools built for SoDaH aim to provide an open-access resource to facilitate synthesis of soil carbon data. Moreover, SoDaH allows for individual locations to contribute results from experimental manipulations, repeated measurements from long-term studies, and local- to regional-scale gradients across ecosystems or landscapes. Finally, we also provide data visualization and analysis tools that can be used to query and analyze the aggregated database. The SoDaH v1.0 dataset is archived and available at https://doi.org/10.6073/pasta/9733f6b6d2ffd12bf126dc36a763e0b4 (Wieder et al., 2020).

54 ENVIRONMENTAL SCIENCES↗

Code for the manuscript "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Mode

We disclose a python/pytorch implementation of the physics-informed machine learning algorithm described in "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Modeling", LA-UR-24-30678. Direct numerical simulation (DNS) of ubiquitous turbulence phenomena is computationally infeasible for realistic flows. As a result, reduced modeling for turbulent flows aim to reduce the number of resolved scales while retaining accurate representations of the small-scale physics. The dynamics of the velocity gradient tensor (VGT) is a key ingredient in reduced or subgrid turbulence models. The evolution equation for the VGT involves nonlocal terms, requiring closure modeling. This implementation of the novel methodology of Lagrangian Attention Tensor Networks (LATN), utilizes a structured representation of the history of the VGT to inform a physics-informed machine learning algorithm. This addition of structured memory terms is shown to outperform previous models when trained and evaluated on DNS data.

Livescu, Daniel [LANL]↗

Sensitivity of magnetic islands in permanent magnet stellarators using the gradient and Hessian methods

Stellarator plasmas are known to be very sensitive to perturbations in the magnetic field. The permanent magnet stellarator was in part developed as a solution to high machining tolerances placed on the shape properties of electromagnetic coils in traditional stellarators. However, as a consequence of this high sensitivity to the field structure, sensitivities of permanent magnet stellarator plasmas to perturbations of permanent magnet properties must necessarily be well-understood. The gradient and Hessian matrix methods have been previously demonstrated to be useful sensitivity analysis methods for modular coils. We apply these two methods to the study of island width sensitivities in both the MUSE and PM4STELL permanent magnet stellarator projects. These sensitivity methods were used to determine the relative impacts of permanent magnet parameter perturbations on island widths in the vacuum field approximation of both stellarator equilibria. The square of resonant magnetic field perturbation is used here as a proxy for island width. In particular, gradients of magnetizations of individual magnets were examined in MUSE, as well as gradients of magnet group displacements informed by device design. Three different forms of permanent magnet magnetization perturbations are investigated for MUSE, and the flux surface response to perturbations is demonstrated. The Hessian matrix method is applied to PM4STELL, illustrating the sensitivity of dominant island widths to displacements of toroidal wedge structures. These methods allow for selective direction of experimental resources toward regions of heightened sensitivity, while constraints on less impactful permanent magnet parameters can be relaxed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A method for convex black-box integer global optimization

Here we study the problem of minimizing a convex function on a nonempty, finite subset of the integer lattice when the function cannot be evaluated at noninteger points. We propose a new underestimator that does not require access to (sub)gradients of the objective; such information is unavailable when the objective is a blackbox function. Rather, our underestimator uses secant linear functions that interpolate the objective function at previously evaluated points. These linear mappings are shown to underestimate the objective in disconnected portions of the domain. Therefore, the union of these conditional cuts provides a nonconvex underestimator of the objective. We propose an algorithm that alternates between updating the underestimator and evaluating the objective function. We prove that the algorithm converges to a global minimum of the objective function on the feasible set. We present two approaches for representing the underestimator and compare their computational effectiveness. We also compare implementations of our algorithm with existing methods for minimizing functions on a subset of the integer lattice. We discuss the difficulty of this problem class and provide insights into why a computational proof of optimality is challenging even for moderate problem sizes.

97 MATHEMATICS AND COMPUTING↗

Event‐Based Training in Label‐Limited Regimes

Abstract The distribution of attributes assigned using data on independent sensors for a specific source, for example, magnitude, can be richly descriptive for final event characterization and associated uncertainty. Attribute distributions can also provide powerful context for event characterization in the absence of comprehensive annotation. This work develops a way to leverage distributional information across a set of sensors in the absence of comprehensive annotation as a domain‐informed regularization term applied during gradient‐based learning. The regularization term is the basis of event‐based training which I show can be a powerful semi‐supervised learning (SSL) approach. I first use a simple feed forward neural network and a toy data set to outline how data set structure interacts with the assumptions inherent to many semi‐supervised learning approaches. I then demonstrate the effectiveness of event‐based training using a deep convolutional neural network for seismic event classification in Utah, which increases SSL accuracy from 92% to 97% on event classification with a limited number of training labels.

Linville, Lisa M.↗

Can Simple Machine Learning Tools Extend and Improve Temperature-Based Methods to Infer Streambed Flux?

Temperature-based methods have been developed to infer 1D vertical exchange flux between a stream and the subsurface. Current analyses rely on fitting physically based analytical and numerical models to temperature time series measured at multiple depths to infer daily average flux. These methods have seen wide use in hydrologic science despite strong simplifying assumptions including a lack of consideration of model structural error or the impacts of multidimensional flow or the impacts of transient streambed hydraulic properties. We performed a “perfect-model experiment” investigation to examine whether regression trees, with and without gradient boosting, can extract sufficient information from model-generated subsurface temperature time series, with and without added measurement error, to infer the corresponding exchange flux time series at the streambed surface. Using model-generated, synthetic data allowed us to assess the basic limitations to the use of machine learning; further examination of real data is only warranted if the method can be shown to perform well under these ideal conditions. We also examined whether the inherent feature importance analyses of tree-based machine learning methods can be used to optimize monitoring networks for exchange flux inference.

54 ENVIRONMENTAL SCIENCES↗

The Information Length Concept Applied to Plasma Turbulence

A methodology to study statistical properties of anomalous transport in fusion plasma is investigated. Three time traces generated by the full-f gyrokinetic code GKNET are analyzed for this purpose. The time traces consist of heat flux as a function of the radial position, which is studied in a novel manner using statistical methods. The simulation data exhibit transport processes with both medium and long correlation length along the radius. A typical example of a phenomenon with long correlation length is avalanches. In order to investigate the evolution of the turbulent state, two basic configurations are studied, one flux-driven and one gradient-driven with decaying turbulence. The information length concept in tandem with Boltzmann–Gibbs and Tsallis entropy is used in the investigation. It is found that the dynamical states in both flux-driven and gradient-driven cases are surprisingly similar, but the Tsallis entropy reveals differences between them. This indicates that the types of probability distribution function are nevertheless quite different since the higher moments are significantly different.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Geothermal Play-Fairway Analysis of Washington State Prospects: Final Report

The Washington State Geothermal Play-Fairway Analysis overcomes the exploration challenges posed by dense vegetation, glacial deposits, and extreme precipitation. The geothermal play-fairways we target are locations where heat, permeability, and saturated porosity are present in sufficient volume to provide adequate heat exchange at depths accessible by modern drilling technology. The three study areas lie along the Cascade Range magmatic arc and are near Mount Baker, Mount St. Helens, and the Wind River Valley. The seven-year project is divided into three phases. In Phase 1 we build on a previous statewide assessment of geothermal resources and develop an initial modeling approach. The results are a series of favorability, uncertainty, and risk maps for three targeted study areas. Based on these initial results, we collect new geologic and geophysical data to further refine our modeling and reduce exploration uncertainty in Phase 2. We improve the modeling method to handle the new data and update the favorability, uncertainty, and risk maps. We also update the conceptual geothermal resource models. In Phase 3 we validate our modeling approach by drilling two temperature-gradient holes and collecting and analyzing core, image logs, and new geochemistry. Our modeling approach improves on an earlier statewide method through a more-rigorous and detailed assessment of heat and permeability. Permeability potential is assessed through geomechanical modeling of the deformation that can generate and maintain reservoir porosity and permeability. Metrics to inform heat potential include temperature-gradient wells, which are sparse in Washington; proximity of Quaternary volcanic vents and young intrusive rock; spring temperature; and reservoir temperature inferred from geothermometry. We weight the individual components using an expert-guided approach known as the Analytical Hierarchy Process. During Phase 2 we also develop a fluid-filled fracture model, and an infrastructure model that helps to delineate areas which are more favorable for geothermal development based on proximity to transmission lines, elevation, land ownership and use restrictions, and availability of process water. New geologic and geophysical data is collected during Phase 2 in each of our three main study areas. At Mount Baker and north of Mount St. Helens we conduct 1:24,000-scale geologic mapping and lidar analysis to better constrain the location and character of surface faults; detailed mapping in the Wind River Valley was completed just prior to the start of this project. Ages of intrusive rocks are determined with 40 Ar/ 39 Ar geochronology, though all of our samples are Miocene or older. We collect ground based gravity observations (a total of 1,580 new stations) in all of our study areas and ground-based magnetic lines (a total of 93 km) at Mount Baker. These data are combined with existing gravity and aeromagnetic data and used to constrain fault locations and geometry. Two to three cross sections are constructed at each study area using the mapped surface geology and forward-modeling of the gravity and magnetic data; these cross sections form the basis for our updated conceptual models. We collect magnetotelluric surveys at Mount Baker and Mount St. Helens and these data are inverted to form a resistivity model from the surface to about 10 km depth; each model shows conductive zones that can be interpreted as upwelling geothermal fluids. At Mount St. Helens we deploy a passive seismic array and use the newly detected events to refine the location of the Saint Helens seismic zone. We also employ ambient-noise tomography to develop a detailed seismic-velocity model for the study area and use this model to help constrain our cross sections and conceptual model. Based on the new data collected during Phase 2—and our updated models—we develop a campaign of temperature-gradient holes and core analysis to validate our modeling in Phase 3. Drill hole MB76-31 is located near Little Park Creek, 11 km west-southwest of the summit of Mount Baker, and is 1,471 ft deep. About 410 ft of core from the lower portion of the hole—and image logs from ~175 ft below ground surface to the bottom—are collected and analyzed. Water samples are collected and processed for geothermometry. Drill hole MSH17-24 is located along upper Schultz Creek, 16 km north-northeast of Mount St. Helens and has core from 470 ft to the bottom at 1,053 ft. We did not collect image logs due to borehole stability concerns, but water samples are collected and analyzed for geothermometry. Repeat temperature-gradient measurements are made at both sites and thermal conductivity is measured from core samples. At MB76-31, the equilibrated temperature gradient of 64°C/km and calculated heat flow of 141–159 mW/m 2 is more than twice the regional average. Detailed mapping and analysis of the core, coupled with correlation to the image logs, indicates a history of permeability generation consistent with our predictions of high permeability. Because the site has high favorability in the Phase 2 model, we consider the results a positive validation of the modeling. At site MSH17-24, the equilibrated temperature gradient of ~15°C/km and calculated heat flow of 41–43 mW/m 2 are similar to regional. Geochemical analysis of the water samples indicates a meteoric source without any geothermal component. Detailed outcrop-based mapping of fault exposures near the drill site and analysis of image logs from nearby boreholes indicates a history of permeability generation consistent with our predictions. Because the site has low favorability in the Phase 2 model, we consider the results a positive validation of the modeling. Together, the two sites provide a reasonably positive validation of the Phase 2 modeling and should encourage future use of this modeling approach.

15 GEOTHERMAL ENERGY↗

An adaptive Hessian approximated stochastic gradient MCMC method

Bayesian approaches have been successfully integrated into training deep neural networks. One popular family is stochastic gradient Markov chain Monte Carlo methods (SG-MCMC), which have gained increasing interest due to their ability to handle large datasets and the potential to avoid overfitting. Although standard SG-MCMC methods have shown great performance in a variety of problems, they may be inefficient when the random variables in the target posterior densities have scale differences or are highly correlated. Here, we present an adaptive Hessian approximated stochastic gradient MCMC method to incorporate local geometric information while sampling from the posterior. The idea is to apply stochastic approximation (SA) to sequentially update a preconditioning matrix at each iteration. The preconditioner possesses second-order information and can guide the random walk of a sampler efficiently. Instead of computing and saving the full Hessian of the log posterior, we use limited memory of the samples and their stochastic gradients to approximate the inverse Hessian-vector multiplication in the updating formula. Moreover, by smoothly optimizing the preconditioning matrix via SA, our proposed algorithm can asymptotically converge to the target distribution with a controllable bias under mild conditions. To reduce the training and testing computational burden, we adopt a magnitude-based weight pruning method to enforce the sparsity of the network. Our method is user-friendly and demonstrates better learning results compared to standard SG-MCMC updating rules. The approximation of inverse Hessian alleviates storage and computational complexities for large dimensional models. Numerical experiments are performed on several problems, including sampling from 2D correlated distribution, synthetic regression problems, and learning the numerical solutions of heterogeneous elliptic PDE. The numerical results demonstrate great improvement in both the convergence rate and accuracy.

97 MATHEMATICS AND COMPUTING↗

Chemical state mapping of simulant Chernobyl lava-like fuel containing material using micro-focused synchrotron X-ray spectroscopy

Uranium speciation and redox behaviour is of critical importance in the nuclear fuel cycle. X-ray absorption near-edge spectroscopy (XANES) is commonly used to probe the oxidation state and speciation of uranium, and other elements, at the macroscopic and microscopic scale, within nuclear materials. Two-dimensional (2D) speciation maps, derived from microfocus X-ray fluorescence and XANES data, provide essential information on the spatial variation and gradients of the oxidation state of redox active elements such as uranium. In the present work, we elaborate and evaluate approaches to the construction of 2D speciation maps, in an effort to maximize sensitivity to the U oxidation state at the U L 3 -edge, applied to a suite of synthetic Chernobyl lava specimens. Our analysis shows that calibration of speciation maps can be improved by determination of the normalized X-ray absorption at excitation energies selected to maximize oxidation state contrast. The maps are calibrated to the normalized absorption of U L 3 XANES spectra of relevant reference compounds, modelled using a combination of arctangent and pseudo-Voigt functions (to represent the photoelectric absorption and multiple-scattering contributions). We validate this approach by microfocus X-ray diffraction and XANES analysis of points of interest, which afford average U oxidation states in excellent agreement with those estimated from the chemical state maps. This simple and easy-to-implement approach is general and transferrable, and will assist in the future analysis of real lava-like fuel-containing materials to understand their environmental degradation, which is a source of radioactive dust production within the Chernobyl shelter.

36 MATERIALS SCIENCE↗

Development of Automated Pipeline for Time-Resolved Link-Wise Vehicular Energy Consumption in the Chattanooga, TN Road Network

The Department of Energy (DOE) has shown strong interest in detecting energy inefficiencies in regional road networks, so as to derive energy consumed at a high spatial temporal resolution. We have developed a workflow to automate the estimation of time-resolved vehicular energy consumption over each link in a road network of interest. The road network used in the current work is centered around the city of Chattanooga, Tennessee and its bordering regions. Utilizing the most mature road network for the Chattanooga, TN region, vehicle speed & count data from TomTom in conjunction with machine learning methods, we have developed an automated pipeline to estimate energy consumption for every link in the network. The first step in the pipeline is ingesting vehicle probe counts and speed estimates from TomTom API. In the next step, the probe counts, speed profiles and other exogenous data (i.e. road types, weather data, ground-truth volume counts and more) were used as input to a supervised learning algorithm to estimate the number of vehicles throughout the entire region for each road segment. These volume estimates were then mapped to a unified road network that contained additional important information such as percentage change in gradient across a link, number of lanes and link lengths that are features in pre-trained single vehicle energy-consumption models available with the RouteE software developed at NREL. The per vehicle energy consumption on each road link predicted using appropriate RouteE vehicular models were multiplied by the volume estimate for the corresponding link over a given time period to predict energy consumed per link for the time interval of interest. Currently, work is underway to improve both the RouteE per vehicle energy estimate and the volume estimates derived from TomTom probe counts. We have also explored the correlation of the link-wise energy estimates with the features of the pre-trained RouteE machine learning model in order to gain insight into what factors contribute most to the link-wise energy consumption.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION,↗

Understanding the fundamental driver of semiconductor radiation tolerance with experiment and theory

Space is the operating environment of a multitude of systems that our society is heavily reliant on today, however, maintaining operability there necessitates special consideration of the electronic systems' tolerance of space radiation. Electronic systems are critically dependent on the electronic properties of their semiconductor components, which are modified by space radiation with an adverse impact on the space system performance. What innate property allows some semiconductors to sustain little damage while others accumulate defects rapidly with dose is poorly understood, which limits the extent to which radiation tolerance can be implemented as a design criterion. Here, to gain insight into what properties are drivers of semiconductor radiation tolerance, the first step is to generate a dataset of the relative radiation tolerance of a broad sampling of semiconductors. To accomplish this, Rutherford backscatter channeling experiments are used to compare the displaced lattice atom buildup in InAs, InP, GaP, GaN, ZnO, MgO, and Si as a function of stepwise alpha particle dose. With this experimental information on radiation-induced incorporation of interstitial defects in hand, hybrid density functional theory electron densities (and their derived quantities) are calculated and their gradient and Laplacian are evaluated to obtain key fundamental information about the interactions in each material. It is shown that simple, undifferentiated values (which are typically used to describe bond strength) are insufficient to predict radiation tolerance. Instead, the curvature of the electron density at bond critical points provides a measure of radiation tolerance consistent with the experimental results obtained. This curvature and associated forces surrounding bond critical points have the potential to disfavor the localization of displaced lattice atoms at these points, favoring their diffusion toward perfect lattice positions. With this criterion to predict radiation tolerance, simple density functional theory simulations can be conducted on potential new materials to gain insight into how they may operate in demanding high radiation environments.

36 MATERIALS SCIENCE↗

Performance of a Natural Gas Solid Oxide Fuel Cell System With and Without Carbon Capture

The fuel cell program at the United States Department of Energy (DOE) National Energy Technology Laboratory (NETL) is focused on the development of low-cost, highly efficient, and reliable fossil-fuel-based solid oxide fuel cell (SOFC) power systems that can generate environmentally-friendly electric power with at least 90 percent carbon capture. NETL’s SOFC technology development roadmap is aligned with near-term market opportunities in the distributed generation sector to validate and advance the technology while paving the way for utility-scale natural gas (NG)- and coal-derived synthesis gas-fueled applications via progressively larger system demonstrations. The present study represents a part of a series of system evaluations being carried out at NETL to aid in prioritizing technological advances along research pathways to the realization of utility-scale SOFC systems, a transformational goal of the fuel cell program. In particular, the system performance of utility-scale NG fuel cell (NGFC) systems with and without carbon dioxide (CO2) capture is presented. The NGFC system analyzed features an external auto-thermal reformer (ATR) feeding the fuel to the SOFC system consisting of planar anode-supported SOFC with separated anode and cathode off-gas streams. In systems with CO2 capture, an air separation unit (ASU) is used to provide the oxygen for the ATR and for the combustion of unutilized fuel in the SOFC anode exhaust along with a CO2 purification unit to provide a nearly pure CO2 stream suitable for transport for usage in enhanced oil recovery operations or for storage in underground saline formations. Remaining thermal energy in the exhaust gases is recovered in a bottoming steam Rankine cycle while supplying any process heat requirements. A reduced order model (ROM) developed at the Pacific Northwest National Laboratory (PNNL) is used to predict the SOFC performance. The ROM, while being computationally effective for system studies, provides other detailed information about the state of the stack, such as the internal temperature gradient, generally not available from simple performance models often used to represent the SOFC. Such additional information can be important in system optimization studies to preclude operation under off-design conditions that can adversely impact overall system reliability. The NGFC system performance was analyzed by varying salient system parameters, including the percent of internal (to the SOFC module) NG reformation—ranging from 0 to 100 percent—fuel utilization, and current density. The impact of advances in underlying SOFC technology on electrical performance was also explored.

solid oxide fuel cell (SOFC), natural gas fuel cel↗

Extending Parsimonious Bayesian Inference

Parsimonious Bayesian inference is a theoretical framework for efficient data assimilation that seeks to balance increased consistency between predictions and training data against corresponding increases in model complexity. Within this framework, over-training is understood as optimization that encodes excessive information within model parameters while only achieving small improvements between predictions and training data. This project aims to develop practical methods of limiting excess model information during optimization. One key observation is that practical heuristics for parsimonious learning in high-dimensions must balance expressivity, i.e. the ability of the model to capture diverse predictions with only a few non-zero parameters, against discoverability, i.e. the ability to train the model with gradient-based optimization and drive parameters to low information states. As such, we developed logical activation functions that are able to adaptively approximate arbitrary truth tables that define Boolean logic operations within a probabilistic framework. These functions have demonstrated the ability to learn exclusive disjunction (XOR) and conditioned disjunction (if [condition] then [result_if_true] else [result_if_false]) within a single layer of a neural network. To efficiently exploit these activation functions to drive parsimonious learning required several other advances within the domain of variational inference. The most efficient form of complexity suppression is structured sparsification, driving most model parameters to zero while achieving the structural coherence among nonzeros needed for bandwidth reduction. Such models are not only far more efficient at suppressing information-theoretic complexity, they also reduce the other forms of complexity (computations, communication, storage, and the number of dependencies needed to evaluate predictions). Aiming to support enhanced sparsification, this project examined new approaches to high-dimensional variational inference that allow us to calibrate and control parameter uncertainty during optimization. By identifying which parameters can sustain sparsifying perturbations with little impact on prediction quality, we can develop better pruning strategies by framing them as approximate Bayesian inference. These advances also open paths to mitigate concerns with deploying advanced learning methods in resource-constrained environments, such as running models on power-limited or communication-limited devices.

97 MATHEMATICS AND COMPUTING↗

A Scalable Gradient Free Method for Bayesian Experimental Design with Implicit Models

Bayesian experimental design (BED) is to answer the question that how to choose designs that maximize the information gathering. For implicit models, where the likelihood is intractable but sampling is possible, conventional BED methods have difficulties in efficiently estimating the posterior distribution and maximizing the mutual information (MI) between data and parameters. Recent work proposed the use of gradient ascent to maximize a lower bound on MI to deal with these issues. However, the approach requires a sampling path to compute the pathwise gradient of the MI lower bound with respect to the design variables, and such a pathwise gradient is usually inaccessible for implicit models. In this paper, we propose a novel approach that leverages recent advances in stochastic approximate gradient ascent incorporated with a smoothed variational MI estimator for efficient and robust BED. Without the necessity of pathwise gradients, our approach allows the design process to be achieved through a unified procedure with an approximate gradient for implicit models. Several experiments show that our approach outperforms baseline methods, and significantly improves the scalability of BED in high-dimensional problems.

Zhang, Jiaxin↗

MFNets: data efficient all-at-once learning of multifidelity surrogates as directed networks of information sources

We present an approach for constructing a surrogate from ensembles of information sources of varying cost and accuracy. The multifidelity surrogate encodes connections between information sources as a directed acyclic graph, and is trained via gradient-based minimization of a nonlinear least squares objective. While the vast majority of state-of-the-art assumes hierarchical connections between information sources, our approach works with flexibly structured information sources that may not admit a strict hierarchy. The formulation has two advantages: (1) increased data efficiency due to parsimonious multifidelity networks that can be tailored to the application; and (2) no constraints on the training data—we can combine noisy, non-nested evaluations of the information sources. Finally, numerical examples ranging from synthetic to physics-based computational mechanics simulations indicate the error in our approach can be orders-of-magnitude smaller, particularly in the low-data regime, than single-fidelity and hierarchical multifidelity approaches.

97 MATHEMATICS AND COMPUTING↗