Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Complex Network Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Characterization and Analysis of the Energy-Reporting Accuracy of Connected Devices

Emerging energy-efficient building systems increasingly exhibit greater functionality, often requiring multiple operating modes (e.g. white-tunability for lighting products and data traffic for devices with networked, integrated sensors). This increased functionality makes energy consumption estimates more complex. Given that these functions consume energy, the energy performance of such building systems is dependent on what operating modes they use and how much time they spend in each mode. Devices and systems that can report their own energy consumption mitigate this energy-performance uncertainty. This study explores the energy-reporting accuracy of market-available connected electrical outlets. The study considers two residential-market products (five units each, one outlet per unit) and three commercial-market products (two units each, 18 to 24 outlets per unit) with the ability to report power drawn and/or energy consumed by devices connected to their receptacles. The products were purchased through typical market channels. Pacific Northwest National Laboratory (PNNL) conducted testing in December 2018 at its Connected Lighting Test Bed (CLTB), using a custom-developed test setup and method adapted from industry standards. The setup collected energy-consumption data reported by the outlet devices under test (DUTs) at one-minute intervals and compared that data with measurements taken by a reference meter over a range of test conditions. The residential products reported power draw but not interval or cumulative energy consumption. The commercial products reported both power draw and cumulative energy consumption. Relative reporting error (RRE) was calculated for all measurements, and analysis of the results revealed variations across devices and test conditions. The total number of measurements (50 for each residential product, 60 for each commercial product) offers an appreciable comparison of performance at the make/model level. The average RRE of the residential products derived from reported power draw was -0.02% and -1.20%. The average RRE for two of the three the commercial products derived from reported power draw was worse than those of the residential products (-2.40%, -2.72%, -0.36%). The internal integration of power over time, used to calculate cumulative energy consumption, typically occurs at current and voltage sampling rates much higher than once per minute. This suggests that the average commercial-product RRE derived from reported energy consumption should be very consistent and better than performance based on reported power draw. However, the RRE derived from reported energy consumption varied significantly across the three makes of commercial-market products and was uniformly less accurate than performance based on reported power draw. Subsequent analysis identified a number of root causes for this decrease in performance, most of which were related to reporting resolution. The goals of this study are to generate awareness of building systems capable of reporting their own energy consumption, further interest in the value of energy data for a variety of uses, draw attention to how the accuracy of reported metrics can be characterized, and quantify the performance variation found in marketavailable products. The results of this study and subsequent related work may be relevant to stakeholders in industry-specification and standards-development organizations. The methods this study employs could inform test and measurement procedures and performance classifications for connected outlets, lighting products, and other building systems capable of reporting their own energy consumption. The study concludes with stakeholder recommendations, including the following: • Energy-reporting device and system manufacturers developing products that report energy consumption should characterize the accuracy of reported metrics using a reference meter calibrated by an independent laboratory that was accredited by an ILAC MRA signatory (and whose scope of accreditation explicitly covers energy measurement), and should include this information on product data sheets. • Standards and specification development organizations should develop application-specific performance classifications that end users can understand and relate to their energy-data use needs (e.g., 2% accuracy class for utility streetlight energy billing needs, or 10% accuracy class for ESCO performance verification needs). • Current or potential owners, operators, and specifiers of energy-reporting building systems should rigorously analyze the dependency of current and planned energy-data use cases on accuracy, noting in particular the dependence (or lack thereof) on relative vs. absolute accuracy, and on trueness vs. precision (i.e., repeatability), and should communicate use-case needs to industry standards and specification organizations.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Characterization and Analysis of the Energy-Reporting Accuracy of Connected Devices

Emerging energy-efficient building systems increasingly exhibit greater functionality, often requiring multiple operating modes (e.g. white-tunability for lighting products and data traffic for devices with networked, integrated sensors). This increased functionality makes energy consumption estimates more complex. Given that these functions consume energy, the energy performance of such building systems is dependent on what operating modes they use and how much time they spend in each mode. Devices and systems that can report their own energy consumption mitigate this energy-performance uncertainty. This study explores the energy-reporting accuracy of market-available connected electrical outlets. The study considers two residential-market products (five units each, one outlet per unit) and three commercial-market products (two units each, 18 to 24 outlets per unit) with the ability to report power drawn and/or energy consumed by devices connected to their receptacles. The products were purchased through typical market channels. Pacific Northwest National Laboratory (PNNL) conducted testing in December 2018 at its Connected Lighting Test Bed (CLTB), using a custom-developed test setup and method adapted from industry standards. The setup collected energy-consumption data reported by the outlet devices under test (DUTs) at one-minute intervals and compared that data with measurements taken by a reference meter over a range of test conditions. The residential products reported power draw but not interval or cumulative energy consumption. The commercial products reported both power draw and cumulative energy consumption. Relative reporting error (RRE) was calculated for all measurements, and analysis of the results revealed variations across devices and test conditions. The total number of measurements (50 for each residential product, 60 for each commercial product) offers an appreciable comparison of performance at the make/model level. The average RRE of the residential products derived from reported power draw was -0.02% and -1.20%. The average RRE for two of the three the commercial products derived from reported power draw was worse than those of the residential products (-2.40%, -2.72%, -0.36%). The internal integration of power over time, used to calculate cumulative energy consumption, typically occurs at current and voltage sampling rates much higher than once per minute. This suggests that the average commercial-product RRE derived from reported energy consumption should be very consistent and better than performance based on reported power draw. However, the RRE derived from reported energy consumption varied significantly across the three makes of commercial-market products and was uniformly less accurate than performance based on reported power draw. Subsequent analysis identified a number of root causes for this decrease in performance, most of which were related to reporting resolution. The goals of this study are to generate awareness of building systems capable of reporting their own energy consumption, further interest in the value of energy data for a variety of uses, draw attention to how the accuracy of reported metrics can be characterized, and quantify the performance variation found in marketavailable products. The results of this study and subsequent related work may be relevant to stakeholders in industry-specification and standards-development organizations. The methods this study employs could inform test and measurement procedures and performance classifications for connected outlets, lighting products, and other building systems capable of reporting their own energy consumption. The study concludes with stakeholder recommendations, including the following: • Energy-reporting device and system manufacturers developing products that report energy consumption should characterize the accuracy of reported metrics using a reference meter calibrated by an independent laboratory that was accredited by an ILAC MRA signatory (and whose scope of accreditation explicitly covers energy measurement), and should include this information on product data sheets. • Standards and specification development organizations should develop application-specific performance classifications that end users can understand and relate to their energy-data use needs (e.g., 2% accuracy class for utility streetlight energy billing needs, or 10% accuracy class for ESCO performance verification needs). • Current or potential owners, operators, and specifiers of energy-reporting building systems should rigorously analyze the dependency of current and planned energy-data use cases on accuracy, noting in particular the dependence (or lack thereof) on relative vs. absolute accuracy, and on trueness vs. precision (i.e., repeatability), and should communicate use-case needs to industry standards and specification organizations.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Discriminative analysis of schizophrenia patients using graph convolutional networks: A combined multimodal MRI and connectomics analysis

Introduction Recent studies in human brain connectomics with multimodal magnetic resonance imaging (MRI) data have widely reported abnormalities in brain structure, function and connectivity associated with schizophrenia (SZ). However, most previous discriminative studies of SZ patients were based on MRI features of brain regions, ignoring the complex relationships within brain networks. Methods We applied a graph convolutional network (GCN) to discriminating SZ patients using the features of brain region and connectivity derived from a combined multimodal MRI and connectomics analysis. Structural magnetic resonance imaging (sMRI) and resting-state functional magnetic resonance imaging (rs-fMRI) data were acquired from 140 SZ patients and 205 normal controls. Eighteen types of brain graphs were constructed for each subject using 3 types of node features, 3 types of edge features, and 2 brain atlases. We investigated the performance of 18 brain graphs and used the TopK pooling layers to highlight salient brain regions (nodes in the graph). Results The GCN model, which used functional connectivity as edge features and multimodal features (sMRI + fMRI) of brain regions as node features, obtained the highest average accuracy of 95.8%, and outperformed other existing classification studies in SZ patients. In the explainability analysis, we reported that the top 10 salient brain regions, predominantly distributed in the prefrontal and occipital cortices, were mainly involved in the systems of emotion and visual processing. Discussion Our findings demonstrated that GCN with a combined multimodal MRI and connectomics analysis can effectively improve the classification of SZ at an individual level, indicating a promising direction for the diagnosis of SZ patients. The code is available at https://github.com/CXY-scut/GCN-SZ.git .

Chen, Xiaoyi↗

Dynamic Boundary Microgrids Under Privatization Considerations

Microgrids have physical, electrical, and logical (data, network, and ownership) boundaries. To power unserved customer loads during an outage, microgrids can extend the traditional operational boundaries. This can become complex when considering microgrid-to-microgrid (M2M) interactions where sensitive information such as competitive microgrid operational data is not shared. This work proposes an optimization method coordinated between microgrid controllers and distribution management systems that limits data sharing. The method involves a competitive bidding strategy that maximizes unserved load coverage while minimizing resource utilization and sensitive operational data sharing among entities. The work is validated on a two-microgrid system with photovoltaic and energy storage systems and curves of load derived from real world residential buildings datasets. Results show that the proposed method, when applied for three distinct use cases of energy storage sufficiency to cover the predefined boundary and/or the expanded boundary, can successfully select and bid the available load coverage.

Starke, Michael [ORNL] (ORCID:0000000221211195)↗

Neural Networks for Nuclear Reactions in MAESTROeX

We demonstrate the use of neural networks to accelerate the reaction steps in the MAESTROeX stellar hydrodynamics code. A traditional MAESTROeX simulation uses a stiff ODE integrator for the reactions; here, we employ a ResNet architecture and describe details relating to the architecture, training, and validation of our networks. Our customized approach includes options for the form of the loss functions, a demonstration that the use of parallel neural networks leads to increased accuracy, and a description of a perturbational approach in the training step that robustifies the model. We test our approach on millimeter-scale flames using a single-step, 3-isotope network describing the first stages of carbon fusion occurring in Type Ia supernovae. We train the neural networks using simulation data from a standard MAESTROeX simulation, and show that the resulting model can be effectively applied to different flame configurations. This work lays the groundwork for more complex networks, and iterative time-integration strategies that can leverage the efficiency of the neural networks.

79 ASTRONOMY AND ASTROPHYSICS↗

Charge-density based convolutional neural networks for stacking fault energy prediction in concentrated alloys

A descriptor-less machine learning (ML) model based only on charge density images extracted from density functional theory (DFT) is developed to predict stacking fault energies (SFE) in concentrated alloys. The model is based on convolutional neural networks (CNNs) as one of the promising ML techniques for dealing with complex images and data. Identification of correct descriptors is a key bottleneck to develop ML models for predicting materials properties. Often, in most ML models, textbook physical descriptors such as atomic radius, valence charge and electronegativity are used as descriptors which have limitations because these properties change in concentrated alloys when multiple elements are mixed to form a solid solution. Here, we illustrate that, within the scope of DFT, the search for descriptors can be circumvented by electronic charge density, which is the backbone of the Kohn-Sham DFT and describes the system completely. The performance of our model is demonstrated by predicting SFE of concentrated alloys with an RMSE and R 2 of 6.18 mJ/m 2 and 0.87, respectively, validating the accuracy of the proposed approach.

36 MATERIALS SCIENCE↗

Hydrological Perspectives on Integrated, Coordinated, Open, Networked (ICON) Science

Hydrologic sciences depend on data monitoring, analyses, and simulations of hydrologic processes to ensure safe, sufficient, and equal water distribution. These hydrologic data come from but are not limited to primary (lab, plot, and field experiments) and secondary sources (remote sensing, UAVs, hydrologic models) that typically follow FAIR Principles (Findable, Accessible, Interoperable, and Reusable: (go-fair.org)). Easy availability of FAIR data has become possible because the hydrology-oriented organizations have pushed the community to increase coordination of the protocols for generating data and sharing model platforms. In addition, networking at all levels has emerged with an invigorated effort to activate community science efforts that complement conventional data collection methods. However, it has become difficult to decipher various complex hydrologic processes with increasing data. Machine learning, a branch of artificial intelligence, provide more accurate and faster alternatives to better understand different hydrological processes. The Integrated, Coordinated, Open, Networked (ICON) framework provides a pathway for water users to include and respect diversity, equity, and inclusivity. In addition, ICONs support the integration of peoples with historically marginalized identities into this professional discipline of water sciences. This article comprises three independent commentaries about the state of ICON principles in hydrology and discusses the opportunities and challenges of adopting them.

(ICON) principles to address↗

Understanding Twinning and Deformation in High Entropy Alloys

A combination of high strength and high ductility has been observed in multi-principal element alloys due to twin formation attributed to low stacking fault energy (SFE). In the pursuit of low SFE alloys, a key bottleneck is the lack of understanding of the composition–SFE cor- relations that would guide tailoring SFE via alloy composition. Using density functional theory (DFT), we show that dopant radius, which have been postulated as a key descriptor for SFE in dilute alloys, does not fully explain SFE trends across different host metals. Instead, charge density is a much more central descriptor. It allows us to (1) explain contrasting SFE trends in Ni and Cu host metals due to various dopants in dilute concentrations, (2) explain the large SFE variations observed in the literature even within a given alloy composition due to the nearest neighbor environments in “model” concentrated alloys, and (3) develop a machine learning model that can be used to predict SFEs in multi-elemental alloys. This model opens a possibility to use charge density as a descriptor for predicting SFE in alloys. Furthermore, a descriptor-less machine learning (ML) model based only on charge density images extracted from density functional theory (DFT) is developed to predict stacking fault energies (SFE) in concentrated alloys. The model is based on convolutional neural networks (CNNs) as one of the promising ML techniques for dealing with complex images and data. Identification of correct descriptors is a key bottleneck to develop ML models for predicting materials properties. Often, in most ML models, textbook physical descriptors such as atomic radius, valence charge and electronegativity are used as descriptors which have limitations because these properties change in concentrated alloys when multiple elements are mixed to form a solid solution. We illustrate that, within the scope of DFT, the search for descriptors can be circumvented by electronic charge density, which is the backbone of the Kohn-Sham DFT and describes the system completely. The performance of our model is demonstrated by predicting SFE of concentrated alloys with an RMSE and R2 of 6.18 mJ/m2 and 0.87, respectively, validating the accuracy of the proposed approach.

36 MATERIALS SCIENCE↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross-dataset tasks. Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross-dataset tasks (up to $28\%$ better relative error and up to almost an order of magnitude better $\chi^2$). Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

Secure Collaborative Environment for Seamless Sharing of Scientific Knowledge

In a secure collaborative environment, tera-bytes of data generated from powerful scientific instruments are used to train secure machine learning (ML) models on exascale computing systems, which are then securely shared with internal or external collaborators as cloud-based services. Devising such a secure platform is necessary for seamless scientific knowledge sharing without compromising individual, or institute-level, intellectual property and privacy details. By enabling new computing opportunities with sensitive data, we envision a secure collaborative environment that will play a significant role in accelerating scientific discovery. Several recent technological advancements have made it possible to realize these capabilities. In this paper, we present our efforts at ORNL toward developing a secure computation platform. We present a use case where scientific data generated from complex instruments, like those at the Spallation Neutron Source (SNS), are used to train a differential privacy enabled deep learning (DL) network on Summit, which is then hosted as a secure multi-party computation (MPC) service on ORNL’s Compute and Data Environment for Science (CADES) cloud computing platform for third-party inference. In this feasibility study, we discuss the challenges involved, elaborate on leveraged technologies, analyze relevant performance results and present the future vision of our work to establish secure collaboration capabilities within and outside of ORNL.

Yoginath, Srikanth↗

Towards replacing physical testing of granular materials with a Topology-based Model

In the study of packed granular materials, the performance of a sample (e.g., the detonation of a high-energy explosive) often correlates to measurements of a fluid flowing through it. The “effective surface area,” the surface area accessible to the airflow, is typically measured using a permeametry apparatus that relates the flow conductance to the permeable surface area via the Carman-Kozeny equation. This equation allows calculating the flow rate of a fluid flowing through the granules packed in the sample for a given pressure drop. However, Carman-Kozeny makes inherent assumptions about tunnel shapes and flow paths that may not accurately hold in situations where the particles possess a wide distribution in shapes, sizes, and aspect ratios, as is true with many powdered systems of technological and commercial interest. To address this challenge, we replicate these measurements virtually on micro-CT images of the powdered material, introducing a new Pore Network Model based on the skeleton of the Morse-Smale complex. Pores are identified as basins of the complex, their incidence encodes adjacency, and the conductivity of the capillary between them is computed from the cross-section at their interface. We build and solve a resistive network to compute an approximate laminar fluid flow through the pore structure. Here, we provide two means of estimating flow-permeable surface area: (i) by direct computation of conductivity, and (ii) by identifying dead-ends in the flow coupled with isosurface extraction and the application of the Carman-Kozeny equation, with the aim of establishing consistency over a range of particle shapes, sizes, porosity levels, and void distribution patterns.

36 MATERIALS SCIENCE↗

Correlator convolutional neural networks as an interpretable architecture for image-like quantum matter data

Image-like data from quantum systems promises to offer greater insight into the physics of correlated quantum matter. However, the traditional framework of condensed matter physics lacks principled approaches for analyzing such data. Machine learning models are a powerful theoretical tool for analyzing image-like data including many-body snapshots from quantum simulators. Recently, they have successfully distinguished between simulated snapshots that are indistinguishable from one and two point correlation functions. Thus far, the complexity of these models has inhibited new physical insights from such approaches. Here, we develop a set of nonlinearities for use in a neural network architecture that discovers features in the data which are directly interpretable in terms of physical observables. Applied to simulated snapshots produced by two candidate theories approximating the doped Fermi-Hubbard model, we uncover that the key distinguishing features are fourth-order spin-charge correlators. Our approach lends itself well to the construction of simple, versatile, end-to-end interpretable architectures, thus paving the way for new physical insights from machine learning studies of experimental and numerical data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Accuracy of predictions made by machine learned models for biocrude yields obtained from hydrothermal liquefaction of organic wastes

Hydrothermal liquefaction (HTL) has potential for converting abundant wet organic wastes into renewable fuels. Because HTL consists of a complex reaction network, deterministic, physics-based prediction of its biocrude yield is prohibitively difficult. Data-driven methods provide an alternative to the physics-based approach; however, rigorous testing must be performed to ensure the accuracy of predictions made by data-driven methods. To this end, a data set was assembled consisting of 570 data points appearing in the open literature. The data set was divided into training, validation, and test sub-sets and used for evaluating different machine learning regression approaches to predict biocrude yield. Among the tested algorithms, Random Forest and eXtreme Gradient Boosting (XGBoost) predicted biocrude yields in a test set that had not been used for training with the greatest accuracy, with root mean square errors (RMSE) of 8.34 and 8.57, respectively. Further refinement of the Random Forest model reduced its RMSE to 8.07. In comparison, predictions of a series of literature models resulted in RMSE ranging from 9.16 in the most accurate case to 27.6 in the least accurate; most literature models yielded RMSE values > 10. Using biocrude yield predictions from the most accurate Random Forest model and a probabilistic economic analysis found that the model accuracy is sufficient to prioritize allocation of resources based on projected minimum fuel selling price. In our report the models and analysis represent a major advance in the ability to use readily available data to predict biocrude yields on new feedstocks that have not previously been studied.

42 ENGINEERING↗

Deep operator network surrogate for phase-field modeling of metal grain growth during solidification

A deep operator network (DeepONet) has been constructed that generates accurate representations of phase-field model simulations for evolving two dimensional metal grain morphology growing from melt. These representations serve as lower resolution, computationally efficient stand-ins for quick parameter space exploration of solutions to the the Allen-Cahn equations that dictate the phase-field model simulations. The experimental target for the phase-field model is a uranium casting system cooling a 434 g uranium charge from a maximum temperature of 1400° C at an average rate of 30° C / min , traversing the crystallographic phases of the pure metal. Experimental parameters inform the phase-field model, whose higher resolution computational model solutions are used to train the DeepONet in a given parameter space with the aim of developing a faster, more efficient method for predicting the solidifying metal's microstructure at different potential experimental values. The final DeepONet generates high accuracy, lower resolution predictions with cumulative relative approximation error over all timesteps of less than 0.5%, while ensuring solutions remain within physically feasible ranges. Further, these relative error values are comparable with other state-of-the-art DeepONet models for microstructure evolution, while significantly reducing the amount of training data required. Training a convolutional neural network simultaneously with the DeepONet, enforcing realistic values at the complex metal grain boundaries, and mathematically encoding boundary conditions into the structure of the DeepONet improved prediction accuracy and computational efficiency over a standard DeepONet model.

36 MATERIALS SCIENCE↗

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Cognition at the Point of Sensing

Over the last 15 years, compressive sensing techniques have been developed which have the potential to greatly reduce the amount of data collected by systems while preserving the amount of information obtained. A cost of this efficiency is that a computationally-intensive optimization routine must be used to put the sensed data into a form that a person can interpret. At the same time, machine learning techniques have experienced tremendous growth as well. Machines have demonstrated the ability learn how to effectively perform tasks such as detection and classification at speeds much faster than humanly possible. Our goal in this project was to study the feasibility of using compressive sensing systems "at the edge." That is, how can compressive sensing sensors be deployed such that information is created at the remote sensor rather than sending raw data to a central processing location? Studies were performed to analyze whether machine learning could be done on the compressively sensed data in its raw form. If a machine is performing the task, is it possible to do so without putting the data into a human interpretable form? We show that this is possible for some systems, in particular a compressive sensing snapshot imaging spectrometer. Machine learning tasks were demonstrated to be more effective and more robust to noise when the machine learning algorithm worked on data in its raw form. This system is shown to outperform a traditional spectrometer. Techniques for reducing the complexity of the reconstruction routine were also analyzed. Techniques for such as data regularization, deep neural networks, and matrix completion were studied and shown to have benefits over traditional reconstruction techniques. In this project we showed that compressive sensing sensors are indeed feasible at the edge. As always, sensors and algorithms must be carefully tuned to work in the constrained environment. In this project we developed tools and techniques to enable those analyses.

47 OTHER INSTRUMENTATION↗

Geometric Measures of Trustworthiness for Machine Learning Predictions

his report details the findings from the research and investigation of Geometric Measures of Trustworthiness for Machine Learning Predictions. We explored the trustworthiness of machine learning (ML) models’ predictions using geometric measures to quantify the similarity of a query point with the training data. Predictive uncertainty in ML can originate from at least three sources: (1) Model uncertainty, which represents the uncertainty in model form (e.g. decision tree, vs neural network) and estimating the model parameters from the training data, (2) Data uncertainty, which represents the natural complexities of the data such as class overlap and inherent noise, and (3) Distributional uncertainty, which represents the mismatch between the training and operational distributions. The proposed measures focus on measuring and explaining the data and distributional uncertainties by measuring the relationships of operational data with the training data.

97 MATHEMATICS AND COMPUTING↗