Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

Multi-Task Learning of Scanning Electron Microscopy and Synthetic Thermal Tomography Images for Detection of Defects in Additively Manufactured Metals

One of the key challenges in laser powder bed fusion (LPBF) additive manufacturing of metals is the appearance of microscopic pores in 3D-printed metallic structures. Quality control in LPBF can be accomplished with non-destructive imaging of the actual 3D-printed structures. Thermal tomography (TT) is a promising non-contact, non-destructive imaging method, which allows for the visualization of subsurface defects in arbitrary-sized metallic structures. However, because imaging is based on heat diffusion, TT images suffer from blurring, which increases with depth. We have been investigating the enhancement of TT imaging capability using machine learning. In this work, we introduce a novel multi-task learning (MTL) approach, which simultaneously performs the classification of synthetic TT images, and segmentation of experimental scanning electron microscopy (SEM) images. Synthetic TT images are obtained from computer simulations of metallic structures with subsurface elliptical-shaped defects, while experimental SEM images are obtained from imaging of LPBF-printed stainless-steel coupons. MTL network is implemented as a shared U-net encoder between the classification and the segmentation tasks. Results of this study show that the MTL network performs better in both the classification of synthetic TT images and the segmentation of SEM images tasks, as compared to the conventional approach when the individual tasks are performed independently of each other.

36 MATERIALS SCIENCE↗

Scalable parallel communications

Coarse-grain parallelism in networking (that is, the use of multiple protocol processors running replicated software sending over several physical channels) can be used to provide gigabit communications for a single application. Since parallel network performance is highly dependent on real issues such as hardware properties (e.g., memory speeds and cache hit rates), operating system overhead (e.g., interrupt handling), and protocol performance (e.g., effect of timeouts), we have performed detailed simulations studies of both a bus-based multiprocessor workstation node (based on the Sun Galaxy MP multiprocessor) and a distributed-memory parallel computer node (based on the Touchstone DELTA) to evaluate the behavior of coarse-grain parallelism. Our results indicate: (1) coarse-grain parallelism can deliver multiple 100 Mbps with currently available hardware platforms and existing networking protocols (such as Transmission Control Protocol/Internet Protocol (TCP/IP) and parallel Fiber Distributed Data Interface (FDDI) rings); (2) scale-up is near linear in n, the number of protocol processors, and channels (for small n and up to a few hundred Mbps); and (3) since these results are based on existing hardware without specialized devices (except perhaps for some simple modifications of the FDDI boards), this is a low cost solution to providing multiple 100 Mbps on current machines. In addition, from both the performance analysis and the properties of these architectures, we conclude: (1) multiple processors providing identical services and the use of space division multiplexing for the physical channels can provide better reliability than monolithic approaches (it also provides graceful degradation and low-cost load balancing); (2) coarse-grain parallelism supports running several transport protocols in parallel to provide different types of service (for example, one TCP handles small messages for many users, other TCP's running in parallel provide high bandwidth service to a single application); and (3) coarse grain parallelism will be able to incorporate many future improvements from related work (e.g., reduced data movement, fast TCP, fine-grain parallelism) also with near linear speed-ups.

Maly, K.↗

The Deep Space Network in the Common Platform Era: A Prototype Implementation at DSS-13

To enhance NASA's Deep Space Network (DSN), an effort is underway to improve network performance and simplify its operation and maintenance. This endeavor, known as the "Common Platform," has both short- and long-term objectives. The long-term work has not begun yet; however, the activity to realize the short-term goals has started. There are three goals for the long-term objective: 1. Convert the DSN into a digital network where signals are digitized at the output of the down converters at the antennas and are distributed via a digital IF switch to the processing platforms. 2. Employ a set of common hardware for signal processing applications, e.g., telemetry, tracking, radio science and Very Long Baseline Interferometry (VLBI). 3. Minimize in-house developments in favor of purchasing commercial off-the-shelf (COTS) equipment. The short-term goal is to develop a prototype of the above at NASA's experimental station known as DSS-13. This station consists of a 34m beam waveguide antenna with cryogenically cooled amplifiers capable of handling deep space research frequencies at S-, X-, and Ka-bands. Without the effort at DSS-13, the implementation of the long-term goal can potentially be risky because embarking on the modification of an operational network without prior preparations can, among other things, result in unwanted service interruptions. Not only are there technical challenges to address, full network implementation of the Common Platform concept includes significant cost uncertainties. Therefore, a limited implementation at DSS-13 will contribute to risk reduction. The benefits of employing common platforms for the DSN are lower cost and improved operations resulting from ease of maintenance and reduced number of spare parts. Increased flexibility for the user is another potential benefit. This paper will present the plans for DSS-13 implementation. It will discuss key issues such as the Common Platform architecture, choice of COTS equipment, and the standard for radio frequency (RF) to digital interface.

Space Communications and Navigation (SCaN)↗

Transportation Network Topologies

A discomforting reality has materialized on the transportation scene: our existing air and ground infrastructures will not scale to meet our nation's 21st century demands and expectations for mobility, commerce, safety, and security. The consequence of inaction is diminished quality of life and economic opportunity in the 21st century. Clearly, new thinking is required for transportation that can scale to meet to the realities of a networked, knowledge-based economy in which the value of time is a new coin of the realm. This paper proposes a framework, or topology, for thinking about the problem of scalability of the system of networks that comprise the aviation system. This framework highlights the role of integrated communication-navigation-surveillance systems in enabling scalability of future air transportation networks. Scalability, in this vein, is a goal of the recently formed Joint Planning and Development Office for the Next Generation Air Transportation System. New foundations for 21PstP thinking about air transportation are underpinned by several technological developments in the traditional aircraft disciplines as well as in communication, navigation, surveillance and information systems. Complexity science and modern network theory give rise to one of the technological developments of importance. Scale-free (i.e., scalable) networks represent a promising concept space for modeling airspace system architectures, and for assessing network performance in terms of scalability, efficiency, robustness, resilience, and other metrics. The paper offers an air transportation system topology as framework for transportation system innovation. Successful outcomes of innovation in air transportation could lay the foundations for new paradigms for aircraft and their operating capabilities, air transportation system architectures, and airspace architectures and procedural concepts. The topology proposed considers air transportation as a system of networks, within which strategies for scalability of the topology may be enabled by technologies and policies. In particular, the effects of scalable ICNS concepts are evaluated within this proposed topology. Alternative business models are appearing on the scene as the old centralized hub-and-spoke model reaches the limits of its scalability. These models include growth of point-to-point scheduled air transportation service (e.g., the RJ phenomenon and the 'Southwest Effect'). Another is a new business model for on-demand, widely distributed, air mobility in jet taxi services. The new businesses forming around this vision are targeting personal air mobility to virtually any of the thousands of origins and destinations throughout suburban, rural, and remote communities and regions. Such advancement in air mobility has many implications for requirements for airports, airspace, and consumers. These new paradigms could support scalable alternatives for the expansion of future air mobility to more consumers in more places.

Holmes, Bruce J.↗

Transportation Network Topologies

A discomforting reality has materialized on the transportation scene: our existing air and ground infrastructures will not scale to meet our nation's 21st century demands and expectations for mobility, commerce, safety, and security. The consequence of inaction is diminished quality of life and economic opportunity in the 21st century. Clearly, new thinking is required for transportation that can scale to meet to the realities of a networked, knowledge-based economy in which the value of time is a new coin of the realm. This paper proposes a framework, or topology, for thinking about the problem of scalability of the system of networks that comprise the aviation system. This framework highlights the role of integrated communication-navigation-surveillance systems in enabling scalability of future air transportation networks. Scalability, in this vein, is a goal of the recently formed Joint Planning and Development Office for the Next Generation Air Transportation System. New foundations for 21st thinking about air transportation are underpinned by several technological developments in the traditional aircraft disciplines as well as in communication, navigation, surveillance and information systems. Complexity science and modern network theory give rise to one of the technological developments of importance. Scale-free (i.e., scalable) networks represent a promising concept space for modeling airspace system architectures, and for assessing network performance in terms of scalability, efficiency, robustness, resilience, and other metrics. The paper offers an air transportation system topology as framework for transportation system innovation. Successful outcomes of innovation in air transportation could lay the foundations for new paradigms for aircraft and their operating capabilities, air transportation system architectures, and airspace architectures and procedural concepts. The topology proposed considers air transportation as a system of networks, within which strategies for scalability of the topology may be enabled by technologies and policies. In particular, the effects of scalable ICNS concepts are evaluated within this proposed topology. Alternative business models are appearing on the scene as the old centralized hub-and-spoke model reaches the limits of its scalability. These models include growth of point-to-point scheduled air transportation service (e.g., the RJ phenomenon and the Southwest Effect). Another is a new business model for on-demand, widely distributed, air mobility in jet taxi services. The new businesses forming around this vision are targeting personal air mobility to virtually any of the thousands of origins and destinations throughout suburban, rural, and remote communities and regions. Such advancement in air mobility has many implications for requirements for airports, airspace, and consumers. These new paradigms could support scalable alternatives for the expansion of future air mobility to more consumers in more places.

Holmes, Bruce J.↗

Deep learning-based segmentation of lithium-ion battery microstructures enhanced by artificially generated electrodes

Accurate 3D representations of lithium-ion battery electrodes, in which the active particles, binder and pore phases are distinguished and labeled, can assist in understanding and ultimately improving battery performance. Here, we demonstrate a methodology for using deep-learning tools to achieve reliable segmentations of volumetric images of electrodes on which standard segmentation approaches fail due to insufficient contrast. We implement the 3D U-Net architecture for segmentation, and, to overcome the limitations of training data obtained experimentally through imaging, we show how synthetic learning data, consisting of realistic artificial electrode structures and their tomographic reconstructions, can be generated and used to enhance network performance. We apply our method to segment x-ray tomographic microscopy images of graphite-silicon composite electrodes and show it is accurate across standard metrics. We then apply it to obtain a statistically meaningful analysis of the microstructural evolution of the carbon-black and binder domain during battery operation.

25 ENERGY STORAGE↗

Network Coordinator Report

This report includes an assessment of the network performance in terms of lost observing time for the 2012 calendar year. Overall, the observing time loss was about 12.3%, which is in-line with previous years. A table of relative incidence of problems with various subsystems is presented. The most significant identified causes of loss were electronics rack problems (accounting for about 21.8% of losses), antenna reliability (18.1%), RFI (11.8%), and receiver problems (11.7%). About 14.2% of the losses occurred for unknown reasons. New antennas are under development in the USA, Germany, and Spain. There are plans for new telescopes in Norway and Sweden. Other activities of the Network Coordinator are summarized.

Himwich, Ed↗

Energy‐aware relay positioning in flying networks

Summary The ability to move and hover has made rotary‐wing unmanned aerial vehicles (UAVs) suitable platforms to act as flying communications relays (FCRs), aiming at providing on‐demand, temporary wireless connectivity when there is no network infrastructure available or a need to reinforce the capacity of existing networks. However, since UAVs rely on their on‐board batteries, which can be drained quickly, they typically need to land frequently for recharging or replacing them, limiting their endurance and the flying network availability. The problem is exacerbated when a single FCR UAV is used. The FCR UAV energy is used for two main tasks: Communications and propulsion. The literature has been focused on optimizing both the flying network performance and energy efficiency from the communications point of view, overlooking the energy spent for the UAV propulsion. Yet, the energy spent for communications is typically negligible when compared with the energy spent for the UAV propulsion. In this article, we propose energy‐aware relay positioning (EREP), an algorithm for positioning the FCR taking into account the energy spent for the UAV propulsion. Building upon the conclusion that hovering is not the most energy‐efficient state, EREP defines the trajectory and speed that minimize the energy spent by the FCR UAV on propulsion, without compromising in practice the quality of service offered by the flying network. The EREP algorithm is evaluated using simulations. The obtained results show gains up to 26% in the FCR UAV endurance for negligible throughput and delay degradation.

Rodrigues, Hugo↗

Parallel and Distributed System Simulation

This exploratory study initiated our research into the software infrastructure necessary to support the modeling and simulation techniques that are most appropriate for the Information Power Grid. Such computational power grids will use high-performance networking to connect hardware, software, instruments, databases, and people into a seamless web that supports a new generation of computation-rich problem solving environments for scientists and engineers. In this context we looked at evaluating the NetSolve software environment for network computing that leverages the potential of such systems while addressing their complexities. NetSolve's main purpose is to enable the creation of complex applications that harness the immense power of the grid, yet are simple to use and easy to deploy. NetSolve uses a modular, client-agent-server architecture to create a system that is very easy to use. Moreover, it is designed to be highly composable in that it readily permits new resources to be added by anyone willing to do so. In these respects NetSolve is to the Grid what the World Wide Web is to the Internet. But like the Web, the design that makes these wonderful features possible can also impose significant limitations on the performance and robustness of a NetSolve system. This project explored the design innovations that push the performance and robustness of the NetSolve paradigm as far as possible without sacrificing the Web-like ease of use and composability that make it so powerful.

Dongarra, Jack↗

Neutrino interaction vertex reconstruction in DUNE with Pandora deep learning

The Pandora Software Development Kit and algorithm libraries perform reconstruction of neutrino interactions in liquid argon time projection chamber detectors. Pandora is the primary event reconstruction software used at the Deep Underground Neutrino Experiment, which will operate four large-scale liquid argon time projection chambers at the far detector site in South Dakota, producing high-resolution images of charged particles emerging from neutrino interactions. While these high-resolution images provide excellent opportunities for physics, the complex topologies require sophisticated pattern recognition capabilities to interpret signals from the detectors as physically meaningful objects that form the inputs to physics analyses. A critical component is the identification of the neutrino interaction vertex. Subsequent reconstruction algorithms use this location to identify the individual primary particles and ensure they each result in a separate reconstructed particle. A new vertex-finding procedure described in this article integrates a U-ResNet neural network performing hit-level classification into the multi-algorithm approach used by Pandora to identify the neutrino interaction vertex. The machine learning solution is seamlessly integrated into a chain of pattern-recognition algorithms. The technique substantially outperforms the previous BDT-based solution, with a more than 20% increase in the efficiency of sub-1 cm vertex reconstruction across all neutrino flavours.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Heat load measurements for the PIP-II pHB650 cryomodule

This study presents a brief overview of the 1st and 2nd phases and an in-depth analysis of the 3rd phase heat load testing performed on the pHB650 (prototype High Beta 650 MHz) cryomodule at PIP2IT (PIP-II Injector Test Facility), with a focus on both the results and the methodological advancements that have improved testing efficiency and accuracy. A key challenge identified in the testing campaign is the higher-than-expected heat loads observed in the first PIP-II (Proton Improvement Plan II) prototype cryomodules (pSSR1 and pHB650) tested at PIP2IT. Elevated heat loads are concerning given the fixed capacity of the PIP-II cryoplant that is currently being installed at Fermilab. However, understanding the sources of these elevated heat loads offers a critical opportunity to implement effective heat load mitigations on upcoming PIP-II cryomodules to stay within the available capacity of the PIP-II cryoplant. The study includes a summary of test results, descriptions of measurement procedures, and key observations on parameters directly and indirectly related to heat load measurements. Direct observations include measured heat loads and the effectiveness of JT heat exchanger under varying conditions, while indirect observation analyze factors such as the temperature distribution on the two-phase pipe and relief piping under varying conditions. Thermal acoustic oscillations (TAO) were identified during testing, which was mitigated by replacing the original G10 stem with a stainless steel stem equipped with wipers for the cryomodule cooldown valve. A major innovation during pHB650 Phase 3 testing was the development of an automated Python script to streamline data acquisition, analysis, and reporting of heat load results. This script automatically retrieved data from ACNET (Accelerator Control Network), performed heat load calculations, and generated detailed reports featuring plots and tables. This advancement significantly reduced manual labor and enhanced the thoroughness of data analysis compared to earlier campaigns. The heat load test reports were promptly uploaded to the electronic logbook shortly after each test, enabling rapid feedback and collaboration between the SRF and cryogenic teams. The heat load measurements included various components: HTTS (high-temperature thermal shield), LTTS (low-temperature thermal shield), 2K isothermal and non-isothermal heat loads. Results were recorded both within the cryomodule and between the bayonet can supply and return. Measurements were conducted under different operating conditions such as "standard", "linac", and "simulated dynamic". Additionally, HTTS and LTTS heat loads were calculated in real time, allowing for the tracking of thermal stability and identification of changes during testing, both in steady-state and transient conditions. The results of this testing campaign not only provide valuable insights into the performance of the pHB650 cryomodule but also highlight best practices and lessons learned that will inform future cryomodule testing at PIP2IT. These include adopting automated tools for data analysis, refining real-time measurement capabilities, and emphasizing detailed pre-test planning. The framework established in this campaign aims to set an improved standard for cryomodule testing and heat load reporting in future cryomodule test campaigns.

Porwisiak, D. [Fermilab; Wroclaw Tech. U.]↗

Active Learning A Neural Network Model For Gold Clusters & Bulk From Sparse First Principles Training Data

Small metal clusters are of fundamental scientific interest and of tremendous significance in catalysis. These nanoscale clusters display diverse geometries and structural motifs depending on the cluster size; a knowledge of this size-dependent structural motifs and their dynamical evolution has been of longstanding interest. Given the high computational cost of first-principles calculations, molecular modeling and atomistic simulations such as molecular dynamics (MD) has proven to be an important complementary tool to aid this understanding. Classical MD typically employ predefined functional forms which limits their ability to capture such complex size-dependent structural and dynamical transformation. Neural Network (NN) based potentials represent flexible alternatives and in principle, well-trained NN potentials can provide high level of flexibility, transferability and accuracy on-par with the reference model used for training. A major challenge, however, is that NN models are interpolative and requires large quantities (similar to 10 4 or greater) of training data to ensure that the model adequately samples the energy landscape both near and far-from-equilibrium. A highly desirable goal is minimize the number of training data, especially if the underlying reference model is first-principles based and hence expensive. In this work, we introduce an active learning (AL) scheme that trains a NN model on-the-fly with minimal amount of first-principles based training data. Our AL workflow is initiated with a sparse training dataset (similar to 1 to 5 data points) and is updated on-the-fly via a Nested Ensemble Monte Carlo scheme that iteratively queries the energy landscape in regions of failure and updates the training pool to improve the network performance. Using a representative system of gold clusters, we demonstrate that our AL workflow can train a NN with similar to 500 total reference calculations. Using an extensive DFT test set of similar to 1100 configurations, we show that our AL-NN is able to accurately predict both the DFT energies and the forces for clusters of a myriad of different sizes. Our NN predictions are within 30 meV/atom and 40 meV/angstrom of the reference DFT calculations. Moreover, our AL-NN model also adequately captures the various size-dependent structural and dynamical properties of gold clusters in excellent agreement with DFT calculations and available experiments. We finally show that our AL-NN model also captures bulk properties reasonably well, even though they were not included in the training data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Building an Integrated Ecosystem of Computational and Observational Facilities to Accelerate Scientific Discovery

Future scientific discoveries will rely on flexible ecosystems that incorporate modern scientific instruments, high performance computing resources, parallel distributed data storage, and performant networks across multiple, independent facilities. In addition to connecting physical resources, such an ecosystem presents many challenges in logistics and accessibility, especially in orchestrating computations and experiments that span across leadership computing systems and experimental instruments. Past efforts have typically been application-specific or limited to interfaces for computing resources. This paper proposes a general framework for integrating computation resources and instrument operations, addressing challenges in code development/execution, data staging and collection, software stack, control mechanisms, resource authorization and governance, and hardware integration. We also describe a demonstration use case wherein a Bayesian optimization algorithm running on an edge computing resource guides a scanning probe microscope to autonomously and intelligently characterize a material sample. This science edge ecosystem framework will provide a blueprint for federating multi-institutional, disparate resources and orchestrating scientific workflows across them to enable next-generation discoveries.

Somnath, Suhas↗

Energy-efficient multimodal mobility networks in transportation digital twins: Strategies and optimization

The study proposes a comprehensive Transportation Mobility (TransitMo) framework covering conceptual design, model formulation, optimization, simulation, and impact analysis of the transportation mobility system. TransitMo is composed of a transportation digital twin developed in Simulation of Urban MObility (SUMO) and an Intelligent Traffic Management and Control Center (ITMCC) that identifies the best ways to improve the movement of people within urban areas using various modes of transportation. This study encompasses advanced modeling techniques, algorithms, and strategic testing to optimize energy efficiency and mobility in a multimodal shared mobility network. TransitMo’s practical applications are exemplified through a city-scaled simulation network in Chattanooga, TN, employing demographic data to analyze historical traffic patterns and forecast future demands. Central to this methodology are three models: the User Preference Model (UP), the Energy Consumption Model (EC), and the System Optimization Model (SO). These models work in concert to iteratively devise the optimal travel incentives and minimize the total system cost in a real-time manner. In conclusion, test results verified that the proposed adaptive incentive program and optimized bus scheduling can improve network performance by increasing public transit ridership.

42 ENGINEERING↗

Identifying common stored product insects using automated deep learning methods

Monitoring stored product insect pests is a common practice for post-harvest management of stored grain and grain-based commodities, which helps ensure product quality from harvest to final consumer. Current methods of sampling and monitoring can be time-consuming, labor-intensive, expensive and require expertise in insect identification. Therefore, this study aims to develop an image-based automated identification system for common stored product insect species using deep-learning methods. Top-down images of the common stored product adult insect species of Rhyzopertha dominica, Cryptolestes ferrugineus, Tribolium castaneum, Sitophilus oryzae, and Oryzaephilus surinamensis were acquired and analyzed. Deep learning-based, state-of-the-art Convolutional Neural Networks (CNN) models (ResNet-50, MobileNet-v2, DarkNet-53, and EfficientNet-b0) were fine-tuned with a transfer learning approach to classify the insect species. All models were able to correctly identify the insect species with at least 96% accuracy and with few misclassifications. One issue with trained CNNs is that they do not explain the reasoning for the classification and are often called a “black box”. Therefore, visualization methods called Gradient-weighted Class Activation Mapping (Grad-CAM) were implemented to explore the black box network. The Grad-CAM uses heat maps to highlight the major image features that the network focused on to make insect species predictions. The Grad-CAM verifies the network's prediction and also helps improve network performance. This study contributes to the overall goal of developing a camera-based system for monitoring stored grain insects. As a result, the developed system would empower warehouse, flour mills, and other food facilities with a tool to quickly and accurately identify insect species in stored product environments and could be implemented as part of a close to real-time monitoring system.

60 APPLIED LIFE SCIENCES↗

Comparison of machine learning techniques to optimize the analysis of plutonium surrogate material via a portable LIBS device

The utilization of machine learning techniques has become commonplace in the analysis of optical emission spectra. These methods are often limited to variants of principal components analysis (PCA), partial-least squares (PLS), and artificial neural networks (ANNs). A plethora of other techniques exist and are well established in the world of data science, yet are seldom investigated for their use in spectroscopic problems. In this study, machine learning techniques were used to analyze optical emission spectra of laser-induced plasma from ceria pellets doped with silicon in order to predict silicon content. Additionally, a boosted regression ensemble model was created, and its predictive accuracy was compared to that of traditional PCA, PLS, and ANN regression models. Boosted regression tree ensembles yielded fits with R-squared (R2) values as high as 0.964 and mean-squared errors of prediction (MSEPs) as low as 0.074, providing the most accurate predictive model. Neural networks performed with slightly lower R2 values and higher MSEPs compared to the ensemble methods, thus indicating susceptibility to overfitting.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predictive understanding of the surface tension and velocity of sound in ionic liquids using machine learning

Knowledge of the physical properties of ionic liquids (ILs), such as the surface tension and speed of sound, is important for both industrial and research applications. Unfortunately, technical challenges and costs limit exhaustive experimental screening efforts of ILs for these critical properties. Previous work has demonstrated that the use of quantum-mechanics-based thermochemical property prediction tools, such as the conductor-like screening model for real solvents, when combined with machine learning (ML) approaches, may provide an alternative pathway to guide the rapid screening and design of ILs for desired physiochemical properties. However, the question of which machine-learning approaches are most appropriate remains. In the present study, we examine how different ML architectures, ranging from tree-based approaches to feed-forward artificial neural networks, perform in generating nonlinear multivariate quantitative structure–property relationship models for the prediction of the temperature- and pressure-dependent surface tension of and speed of sound in ILs over a wide range of surface tensions (16.9–76.2 mN/m) and speeds of sound (1009.7–1992 m/s). The ML models are further interrogated using the powerful interpretation method, shapley additive explanations. We find that several different ML models provide high accuracy, according to traditional statistical metrics. The decision tree-based approaches appear to be the most accurate and precise, with extreme gradient-boosting trees and gradient-boosting trees being the best performers. However, our results also indicate that the promise of using machine-learning to gain deep insights into the underlying physics driving structure–property relationships in ILs may still be somewhat premature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗