Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Kimberlina 1.2 CCUS Geophysical Models and Synthetic Data Sets

This synthetic multi-scale and multi-physics data set was produced in collaboration with teams at the Lawrence Berkeley National Laboratory, National Energy Technology Laboratory, Los Alamos National Laboratory, and Colorado School of Mines through the Science-informed Machine Learning for Accelerating Real-Time Decisions in Subsurface Applications (SMART) Initiative. Data are associated with the following publication: Alumbaugh, D., Gasperikova, E., Crandall, D., Commer, M., Feng, S., Harbert, W., Li, Y., Lin, Y., and Samarasinghe, S., “The Kimberlina Synthetic Geophysical Model and Data Set for CO2 Monitoring Investigations”, The Geoscience Data Journal, 2023, DOI: 10.1002/gdj3.191. The dataset uses the Kimberlina 1.2 CO2 reservoir flow model simulations based on a hypothetical CO2 storage site in California (Birkholzer et al., 2011; Wainwright et al., 2013). Geophysical properties models (P- and S-wave seismic velocities, saturated density, and electrical resistivity) were produced with an approach similar to that of Yang et al. (2019) and Gasperikova et al. (2022) for 100 Kimberlina 1.2 reservoir models. Links to individual resources are provided below: [CO2 Saturation Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-co2-saturation-models); Resistivity Models – [part 1](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-1), [part 2](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-2), and [part 3](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-3); [Vp Velocity Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-vp-velocity-models); [Vs Velocity Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-vs-velocity-models); [Density Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-density-models). The 3D distributions of geophysical properties for the 33 time stamps of the SIM001 model were used to generate synthetic seismic, gravity, and electromagnetic (EM) responses for 33 times between zero and 200 years. Synthetic surface seismic data were generated using 2D and 3D finite-difference codes that simulate the acoustic wave equation (Moczo et al., 2007). 2D data were simulated for six point-pressure sources along a 2D line with 10 m receiver spacing and a time spacing of 0.0005 s. 3D simulations were completed for 25 surface pressure sources using a source separation of 1 km in both the x and y directions and a time spacing of 0.001 s. Links to individual resources are provided below: [2D velocity models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-2d-velocity-models) and [2D surface seismic data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-2d-surface-seismic-data). [3D velocity models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-velocity-models), and 3D seismic data [year0](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year0), [year1](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year1), [year2](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year2), [year5](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year5), [year10](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year10), [year15](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year15), [year20](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year20), [year25](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year25), [year30](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year30), [year35](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year35), [year40](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year40), [year45](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year45), [year49](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year49), [year50](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year50), [year51](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year51), [year52](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year52), [year55](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year55), [year60](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year60), [year65](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year65), [year70](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year70), [year75](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year75), [year80](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year80), [year85](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year85), [year90](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year90), [year95](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year95), [year100](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year100), [year110](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year110), [year120](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year120), [year130](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year130), [year140](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year140), [year150](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year150), [year175](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year175), [year200](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year200). The Python scripts to read these models and data are provided [here](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-python-scripts). EM simulations used a borehole-to-surface survey configuration, with the source located near the reservoir level and receivers on the surface using the code developed by Commer and Newman (2008). Pseudo-2D data for the source at [2500 m](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-pseudo-2d-csem-data-tz2500m) and [3025 m](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-pseudo-2d-csem-data-tz3025m), used a 2D inline receiver configuration to simulate a response over 3D resistivity models. The [3D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-csem-data) contain electric fields generated by borehole sources at monitoring well locations and measured over a surface receiver grid. Vector gravity data, both on the surface and in boreholes, were simulated using a modeling code developed by Rim and Li (2015). The simulation scenarios were parallel to those used for the EM: [pseudo-2D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-gravity-data) were calculated along the same lines and within the same boreholes, and [3D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-gravity-data) were simulated over 3D models on the surface and in three monitoring wells. A series of [synthetic well logs](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-well-logs) of CO2 saturation, acoustic velocity, density, and induction resistivity in the injection well and three monitoring wells are also provided at 0, 1, 2, 5, 10, 15, and 20 years after the initiation of injection. These were constructed by combining the low-frequency trend of the geophysical models with the high-frequency variations of actual well logs collected in the Kimberlina 1 well that was drilled at the proposed site. Measurements of permeability and pore connectivity were made on cores of Vedder Sandstone, which forms the primary reservoir unit: [CT micro scans](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-ct-micro-scans-of-vedder-formation) and [Industrial CT Images](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-industrial-ct-images-vedder-formation). These measurements provide the range of scales in the otherwise synthetic data set to be as close to a real-world situation as possible. References: Birkholzer, J.T., Zhou, Q., Cortis, A. and Finsterle, S., 2011. A sensitivity study on regional pressure buildup from large-scale CO2 storage projects. Energy Procedia, 4, 4371-4378. Commer, M., and Newman, G.A., 2008. New advances in three-dimensional controlled-source electromagnetic inversion, Geophysical Journal International, 172, 513-535. Gasperikova, E., Appriou, D., Bonneville, A., Feng, Z., Huang, L., Gao, K., Yang, X., Daley, T., 2022, Sensitivity of geophysical techniques for monitoring secondary CO2 storage plumes, Int. J. Greenh. Gas Control, Volume 114, 103585, ISSN 1750-5836, https://doi.org/10.1016/j.ijggc.2022.103585. Moczo, P., J.O. Robertsson and L. Eisner, 2007, The finite-difference time-domain method for modeling of seismic wave propagation: Advances in geophysics, 48, 421-516. Rim, H., and Y. Li, 2015, Advantages of borehole vector gravity in density imaging, Geophysics, 80, G1-G13. Wainwright, H. M.; Finsterle, S.; Zhou, Q.; Birkholzer, J. T., 2013. Modeling the Performance of Large-Scale CO2 Storage Systems: A Comparison of Different Sensitivity Analysis Methods. International Journal of Greenhouse Gas Control, 17, 189205. https://doi.org/10.1016/j.ijggc.2013.05.007, DOI: 10.18141/1603331. Yang, X., Buscheck, T.A., Mansoor, K., Wang, Z., Gao, K., Huang, L., Appriou, D., and Carroll, S.A., 2019. Assessment of geophysical monitoring methods for detection of brine and CO2 leakage in drinking water aquifers, International Journal of Greenhouse Gas Control, 90, 102803, https://doi.org/10.1016/j.ijggc.2019.102803.

CCUS↗

Opportunities and Challenges from Artificial Intelligence and Machine Learning for the Advancement of Science, Technology, and the Office of Science Missions

In February 2019, the President signed Executive Order 13859, Maintaining American Leadership in Artificial Intelligence. This order launched the American Artificial Intelligence Initiative, a concerted effort to promote and protect AI technology and innovation in the United States. The Initiative implements a government-wide strategy in collaboration and engagement with the private sector, academia, the public, and like-minded international partners. Among other actions, key directives in the Initiative called for Federal agencies to: Prioritize AI research and development investments, Enhance access to high-quality cyberinfrastructure and data, Ensure that the US maintains an international leadership role in the development of technical standards for AI, and Provide education and training opportunities to prepare the American workforce for the new era of AI. The mission of the Department of Energy (DOE) is to ensure America’s security and prosperity by addressing its energy, environmental, and nuclear challenges through transformative science and technology solutions. In terms of Science and Innovation, the DOE’s mission is to maintain a vibrant US effort in science and engineering as a cornerstone of our economic prosperity with clear leadership in strategic areas. From July to October in 2019, the Argonne, Oak Ridge, and Berkeley National Laboratories hosted a series of four AI for Science Town Hall meetings in Chicago, Oak Ridge, Berkeley, and Washington DC. The four meetings were attended by over 1300 scientists from the 17 DOE Labs, 39 companies, and over 90 universities. The goal of the Town Hall series was ‘to examine scientific opportunities in the areas of artificial intelligence, Big Data, and high-performance computing (HPC) in the next decade, and to capture the big ideas, grand challenges, and next steps to realizing these.’ The discussions at the meetings were captured in the final report of the AI for Science Town Hall meetings.

42 ENGINEERING↗

HFTS-1 Natural Joint Data and Engineering Summary

This report provides an analysis of engineering and geologic data collected from the Hydraulic Fracturing Test Site #1 (HTSF-1) project in the southern Midland Basin, Reagan County, Texas. The site is being studied as part of the Science-informed Machine Learning for Accelerating Real-Time Decisions in Subsurface Applications (SMART) Initiative at the National Energy Technology Laboratory (NETL). The data collected is intended to provide a basis of understanding of the site and to construct simulation models. This report provides a summary analysis of the engineering and geologic data collected from the project to provide a basis of understanding of the site and to construct simulation models. The report examines various possible correlations in engineering properties based on natural fracture data from the program, which consisted of 11 horizontal wells, one vertical well, and one slant well. The horizontal wells are in two horizons: 1) Upper and 2) Middle Wolfcamp formations. Program data include: fracture frequency and fracture orientation data from four core runs in a slant well; fracture spacing and orientation from a vertical pilot well; laboratory triaxial testing and mineralogical determinations; and porosity results from magnetic resonance analyses from various wells, together with observations based on the data collection. In addition, available references were reviewed on the site for additional insights. As the focus of the report is on the natural system, hydraulic fracture data from the site were not examined in detail in this report. Data variability is the chief observation in examination of the database. Fracture frequency in the slant well can range from sections with values as high as five fractures per ft to sections up to 100+ ft in length with no natural observed fractures. Fracture spacing across is typically less than 10 ft, but can range up to hundreds of feet. Apparent fracturing shows the trends in two predominate orientations, E-SW and WNW-ESE, but minor variations exist. The rock units vary across the site from siliceous mudstones to calcareous mudstones, showing a general layering with depth. The laboratory properties such as strength and modulus show no apparent trend with depth, but appear to correlate with rock mineralogy with high strength and modulus values where calcium content is high. In addition, an attempt to examine variability and mineralogy on a larger scale was made using a color-coded system based on gamma ray measurements. A staged colored approach was adopted, presuming that lower gamma ray values indicate higher value of calcium content (blue scale) and that higher gamma ray values indicate higher clay mineral content (orange scale). As provided in report appendices, the system correlated well with visual examination of the slant core and the petrofabric analyses of the vertical pilot well. The results showed large variability in mineral content along the horizontal plane across the site. The change in mineral content was also rapid, on a scale less than that of the average hydraulic fracture stage length of about 180 ft.

42 ENGINEERING↗

Overview of SMART Initiative

The objective of the SMART Initiative, i.e., Science-informed Machine Learning (ML) for Accelerating Real-Time Decisions in Subsurface Applications, is to show how the utilization of ML can significantly improve efficiency and effectiveness of field-scale commercial carbon storage operations in three main areas: real-time visualization, virtual learning, and real-time forecasting. This presentation reports the status of SMART initiative for demonstrating: (a) virtual learning during the pre-injection permitting phase, and (b) ML-assisted operational decision making and visualization.

Siriwardane, Hema↗

Deep nonparametric estimation of operators between infinite dimensional spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification via Stable Distribution Propagation

We propose a new approach for propagating stable probability distributions through neural networks. Our method is based on local linearization, which we show to be an optimal approximation in terms of total variation distance for the ReLU non-linearity. This allows propagating Gaussian and Cauchy input uncertainties through neural networks to quantify their output uncertainties. To demonstrate the utility of propagating distributions, we apply the proposed method to predicting calibrated confidence intervals and selective prediction on out-of-distribution data. The results demonstrate a broad applicability of propagating distributions and show the advantages of our method over other approaches such as moment matching.

Artificial Intelligence (cs.AI)↗

Global Framework for Emulation of Nuclear Calculations

We introduce a hierarchical framework that combines ab initio many-body calculations with a Bayesian neural network, developing emulators capable of accurately predicting nuclear properties across isotopic chains simultaneously and being applicable to different regions of the nuclear chart. We benchmark our developments using the oxygen isotopic chain, achieving accurate results for ground-state energies and nuclear charge radii, while providing robust uncertainty quantification. Our framework enables global sensitivity analysis of nuclear binding energies and charge radii with respect to the low-energy constants that describe the nuclear force.

FOS: Computer and information sciences↗

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark -- A Bayesian Inverse UQ-based Approach for Data Assimilation

The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. When comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that GLLS predictions fail to replicate computed response distributions for nonlinear applications, while MOCABA shows near agreement, and IUQ uses computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

FOS: Computer and information sciences↗

Domain Knowledge Guided Bayesian Optimization For Autonomous Alignment Of Complex Scientific Instruments

Bayesian Optimization (BO) is a powerful tool for optimizing complex non-linear systems. However, its performance degrades in high-dimensional problems with tightly coupled parameters and highly asymmetric objective landscapes, where rewards are sparse. In such needle-in-a-haystack scenarios, even advanced methods like trust-region BO (TurBO) often lead to unsatisfactory results. We propose a domain knowledge guided Bayesian Optimization approach, which leverages physical insight to fundamentally simplify the search problem by transforming coordinates to decouple input features and align the active subspaces with the primary search axes. We demonstrate this approach's efficacy on a challenging 12-dimensional, 6-crystal Split-and-Delay optical system, where conventional approaches, including standard BO, TuRBO and multi-objective BO, consistently led to unsatisfactory results. When combined with an reverse annealing exploration strategy, this approach reliably converges to the global optimum. The coordinate transformation itself is the key to this success, significantly accelerating the search by aligning input co-ordinate axes with the problem's active subspaces. As increasingly complex scientific instruments, from large telescopes to new spectrometers at X-ray Free Electron Lasers are deployed, the demand for robust high-dimensional optimization grows. Our results demonstrate a generalizable paradigm: leveraging physical insight to transform high-dimensional, coupled optimization problems into simpler representations can enable rapid and robust automated tuning for consistent high performance while still retaining current optimization algorithms.

FOS: Computer and information sciences↗

Spatial patterns of snow distribution in the sub-Arctic

Abstract. The spatial distribution of snow plays a vital role in sub-Arctic and Arctic climate, hydrology, and ecology due to its fundamental influence on the water balance, thermal regimes, vegetation, and carbon flux. However, the spatial distribution of snow is not well understood, and therefore, it is not well modeled, which can lead to substantial uncertainties in snow cover representations. To capture key hydro-ecological controls on snow spatial distribution, we carried out intensive field studies over multiple years for two small (2017–2019; ∼ 2.5 km2) sub-Arctic study sites located on the Seward Peninsula of Alaska. Using an intensive suite of field observations (> 22 000 data points), we developed simple models of the spatial distribution of snow water equivalent (SWE) using factors such as topographic characteristics, vegetation characteristics based on greenness (normalized different vegetation index, NDVI), and a simple metric for approximating winds. The most successful model was random forest, using both study sites and all years, which was able to accurately capture the complexity and variability of snow characteristics across the sites. Approximately 86 % of the SWE distribution could be accounted for, on average, by the random forest model at the study sites. Factors that impacted year-to-year snow distribution included NDVI, elevation, and a metric to represent coarse microtopography (topographic position index, TPI), while slope, wind, and fine microtopography factors were less important. The characterization of the SWE spatial distribution patterns will be used to validate and improve snow distribution modeling in the Department of Energy's Earth system model and for improved understanding of hydrology, topography, and vegetation dynamics in the sub-Arctic and Arctic regions of the globe.

54 ENVIRONMENTAL SCIENCES↗

Opportunities for Process Intensification with Membranes to Promote Circular Economy Development for Critical Minerals

Critical minerals are essential to the future of clean energy, especially energy storage, electric vehicles, and advanced electronics. In this paper, we argue that process systems engineering (PSE) paradigms provide essential frameworks for enhancing the sustainability and efficiency of critical mineral processing pathways. As a concrete example, we review challenges and opportu-nities across material-to-infrastructure scales for process intensification (PI) with membranes. Within critical mineral processing, there is a need to reduce environmental impact, especially con-cerning chemical reagent usage. Feed concentrations and product demand variability require flex-ible, intensified processes. Further, unique feedstocks require unique processes (i.e., no one-size-fits-all recycling or refining system exists). Membrane materials span a vast design space that allows significant optimization. Therefore, there is a need to rapidly identify the best opportunities for membrane implementation, thus informing materials optimization with process and infrastructure scale performance targets. Finally, scale-up must be accelerated and de-risked across the materials-to-process levels to fully realize the opportunity presented by membranes, thereby fostering the development of a circular economy for critical minerals. Tackling these challenges requires integrating efforts across diverse disciplines. We advocate for a holistic molecular-to-systems perspective for fully realizing PI with membranes to address sustainability challenges in critical mineral processing. The opportunities for PI with membranes are excellent applications for emerging research in machine learning, data science, automation, and optimization.

Dougher, Molly↗

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system and the aviation industry has experienced a steady decrease in fatalities over the years. This can be attributed to both improved flight critical systems with redundant hardware and software protections, as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main approach for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave within the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety, creating labels for the data requires huge amount of effort and is largely impractical. To address this challenge, we developed a Convolutional Variational Auto-Encoder (CVAE), which is an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach as well as unsupervised clustering-based approach using KMeans++ and kernel-based approach using One-Class Support Vector Machine (OC-SVM) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Memarzadeh, Milad↗

Unsupervised Anomaly Detection in High-Dimensional Flight Data Using Convolutional Variational Auto-Encoder

The modern National Airspace System (NAS) is an extremely safe system. The industry has experienced a steady decrease in fatalities over the years. This can be contributed to both improved flight critical systems with redundant hardware and software protections as well as an increased focus on active monitoring and response to real time and historically identified vulnerabilities by implementing more resilient procedures and protocols. The main practice for identifying vulnerabilities in operations leverages domain expertise using knowledge about how the system should behave with the expected tolerances to known safety margins. This approach works well when the system has a well-defined operating condition. However, the operations in the NAS can be highly complex with various nuances that render it difficult to clearly pre-define all known safety vulnerabilities. With the advancement of data science and machine learning techniques, the potential to automatically identify emerging vulnerabilities in the observed operations has become more practical in recent years. The state-of-the-art anomaly detection approaches in aerospace data usually rely on supervised or semi-supervised learning. However, in many real-world problems such as flight safety creating labels for the data requires huge amount of efforts and is largely expensive. As a result, in this article, we develop a Convolutional Variational Auto-Encoder (CVAE), an unsupervised learning approach for anomaly detection in high-dimensional heterogeneous time-series data. We validate performance of CVAE compared to the state-of-the-art supervised learning approach (as an upper bound) as well as an supervised clustering based on K-Means (as a lower bound) on Yahoo!'s benchmark time series anomaly detection data. Finally, we showcase performance of CVAE on a case study of identifying anomalies in the first 60 seconds of commercial flights' take-offs using Flight Operational Quality Assurance (FOQA) data.

Milad Memarzadeh↗