Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

The Intrinsic Scatter of the Broad Lines–Narrow Line Correlation in Type I AGN

A correlation between the FWHM of the broad Balmer lines and the narrow [O iii]/H{sub β} line ratio was recently applied to the black hole (BH) mass estimation in type II active galactic nuclei (AGN), where only the narrow lines are visible to the observer. The correlation was initially quantified with type I AGN using stacked spectra, in groups automatically created using a machine-learning algorithm. Such an analysis does not provide information about the intrinsic scatter of the correlation. In addition, it does not necessarily reproduce the true underlying correlation, for example, due to the stacking of spectra with different properties. Testing these two issues requires measurements of individual objects. In this work, we perform such a test by fitting the broad and narrow lines for 8302 type I AGN from the Sloan Digital Sky Survey. Due to the difficulty in reliably measuring the narrow Balmer lines in such objects, which are, in many cases, a small contribution on top of the broad lines, we visually inspect all of the fits and identify 1561 objects with robust measurements. Using these measurements, we find that while a correlation does exist, it shows a large scatter and is not well described by a linear relation. This should be taken into account when using the broad H{sub β} FWHM versus narrow [O iii]/H{sub β} correlation for type II AGN BH mass estimation.

79 ASTRONOMY AND ASTROPHYSICS↗

Genomic factors shaping codon usage across the Saccharomycotina subphylum

Codon usage bias, or the unequal use of synonymous codons, is observed across genes, genomes, and between species. It has been implicated in many cellular functions, such as translation dynamics and transcript stability, but can also be shaped by neutral forces. We characterized codon usage across 1,154 strains from 1,051 species from the fungal subphylum Saccharomycotina to gain insight into the biases, molecular mechanisms, evolution, and genomic features contributing to codon usage patterns. We found a general preference for A/T-ending codons and correlations between codon usage bias, GC content, and tRNA-ome size. Codon usage bias is distinct between the 12 orders to such a degree that yeasts can be classified with an accuracy >90% using a machine learning algorithm. We also characterized the degree to which codon usage bias is impacted by translational selection. We found it was influenced by a combination of features, including the number of coding sequences, BUSCO count, and genome length. Our analysis also revealed an extreme bias in codon usage in the Saccharomycodales associated with a lack of predicted arginine tRNAs that decode CGN codons, leaving only the AGN codons to encode arginine. Analysis of Saccharomycodales gene expression, tRNA sequences, and codon evolution suggests that avoidance of the CGN codons is associated with a decline in arginine tRNA function. Consistent with previous findings, codon usage bias within the Saccharomycotina is shaped by genomic features and GC bias. However, we find cases of extreme codon usage preference and avoidance along yeast lineages, suggesting additional forces may be shaping the evolution of specific codons.

59 BASIC BIOLOGICAL SCIENCES↗

Discovering Hidden Geothermal Signatures using Unsupervised Machine Learning

Discovering hidden geothermal resources is a very challenging task. It requires the mining of large datasets, including various diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used Play Fairway Analysis (PFA) typically relies on subject-matter expertise to analyze site or regional data to estimate geothermal conditions and prospectivity. Here, we demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset of Southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. However, most of these systems are poorly characterized because of insufficient existing data and limited past explorative studies. This study aims to discover hidden patterns and relationships in the SWNM geothermal dataset to better understand regional hydrothermal conditions. This is achieved by applying an unsupervised machine learning algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden (latent) signatures characterizing datasets, (2) the optimal number of these signatures, (3) dominant data attributes associated with each signature, and (4) spatial distribution of the extracted signatures. Here, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, geothermal attributes at 44 locations in SWNM. NMFk successfully finds data patterns and identifies the spatial associations of hydrothermal signatures with the four physiographic provinces in SWNM (Colorado Plateau, Volcanic Field, Basin and Range, and the Rio Grande rift). The algorithm identified up to 5 hydrothermal signatures in the SWNM datasets that differentiate between low- and medium-temperature hydrothermal systems in different provinces. Also, the algorithm identifies two medium-temperature hydrothermal systems in SWNM that require further exploration for geothermal resource development. Based on our analyses, 12 of the attributes are important to identify medium-temperature hydrothermal systems, and the remaining six attributes are critical to characterize low-temperature hydrothermal systems. Based on the obtained results, we identify potential physiographic provinces for further exploration to characterize them as geothermal resources. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored areas.

58 GEOSCIENCES↗

Deep learning based x-ray spectrometer for high repetition rate characterization of betatron radiation

Betatron radiation produced from a laser-wakefield accelerator is a broadband, hard x-ray (>1 keV) source that has been used in a variety of applications in medicine, engineering, and fundamental science. Further development and optimization of stable, high repetition rate (HRR) (>1 Hz) betatron sources will provide a means to extend their application base to include single-shot dynamical measurements of ultrafast processes or dense materials. Recent advances in laser technology used in such experiments have enabled increases in shot-rate and system stability, providing improved statistical analysis and detailed parameter scans. However, unique challenges exist at high repetition rate, where data throughput and source optimization are now limited by diagnostic acquisition rates and analysis. Here, we present the development of a machine-learning algorithm for the real-time analysis of betatron radiation. We report on the fielding of this deep learning algorithm for online source characterization at the Institut National de la Recherche Scientifique's Advanced Laser Light Source. By fine-tuning an algorithm originally trained on a fully synthetic dataset using a subset of experimental data, the algorithm can predict the betatron critical energy with a percent error of 7.2 % with a reconstruction time of 1.5 ms, providing a valuable tool for real-time, multi-objective optimization at HRR.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Medium Energy Electron Flux in Earth's Outer Radiation Belt (MERLIN): A Machine Learning Model

The radiation belts of the Earth, filled with energetic electrons, comprise complex and dynamic systems that pose a significant threat to satellite operation. While various models of electron flux both for low and relativistic energies have been developed, the behavior of medium energy (120–600 keV) electrons, especially in the MEO region, remains poorly quantified. At these energies, electrons are driven by both convective and diffusive transport, and their prediction usually requires sophisticated 4D modeling codes. In this paper, we present an alternative approach using the Light Gradient Boosting (LightGBM) machine learning algorithm. The Medium Energy electRon fLux In Earth's outer radiatioN belt (MERLIN) model takes as input the satellite position, a combination of geomagnetic indices and solar wind parameters including the time history of velocity, and does not use persistence. MERLIN is trained on >15 years of the GPS electron flux data and tested on more than 1.5 years of measurements. Tenfold cross validation yields that the model predicts the MEO radiation environment well, both in terms of dynamics and amplitudes o f flux. Evaluation on the test set shows high correlation between the predicted and observed electron flux (0.8) and low values of absolute error. The MERLIN model can have wide space weather applications, providing information for the scientific community in the form of radiation belts reconstructions, as well as industry for satellite mission design, nowcast of the MEO environment, and surface charging analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Evaluation of antifouling surfaces using a method that employs mussel larvae settlement quantified by machine learning

Antifouling coating development requires extensive performance testing. Coatings that prevent aquatic larval settlement are of interest because many forms of macrofouling begin at the larval stage. However, field testing can be time consuming and poorly controlled. Herein is reported a screening tool, Settlement of Larvae Assay using Mussels (SLAM), for down-selecting materials prior to field testing. The method entails using a dense concentration of mussel larvae that are allowed to settle on submerged test surfaces. Settled larvae are then quantified to provide a measure of antifouling performance. The SLAM test differentiated coatings with only slight differences in formulation. To enable efficient quantification of dense larvae settlement, an automated counting method was developed that combines two analyses: a color thresholding identifies larvae clumps, and a machine learning algorithm identifies non-clumped larvae. Finally, this automated ‘hybrid’ approach rapidly quantifies settled larvae as effectively as manual counting but in a fraction of the time.

Mytilus↗

Enabling Computation on Sensitive Data in International Safeguards with Privacy-Preserving Encryption Techniques

Privacy-preserving machine learning is a field of study that explores how to protect and preserve the privacy of sensitive data while allowing the data to be used by machine learning algorithms. This field has had substantial industry investment due to heightened concerns about privacy in the technology industry, with a focus in two broad application areas: financial services and healthcare. Numerous privacy-preserving methods have also been proposed for international safeguards, but they have been difficult to enact because the data they require is con- sidered sensitive or proprietary by the nuclear facility operator. This work examines how current privacy-preserving approaches might be used to enable the International Atomic Energy Agency (IAEA) to use that data to contribute to a safeguards conclusion about a state while giving nuclear operators confidence that their sensitive data is adequately protected. This paper begins by exploring several broad categories of privacy-preserving techniques including homomorphic encryption, secure multiparty computation, secure enclaves, and zero-knowledge proofs. Then we discuss some of the security considerations related to using these methods, potential use cases, and a conceptual system design for applying privacy-preserving methods in international safeguards.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

RU Net for Automatic Characterization of TRISO Fuel Cross Sections

TRistructural ISOtropic (TRISO) particle fuel is a type of nuclear fuel known for its high-temperature and high-burnup performance. Each sub-millimeter diameter TRISO particle consists of uranium-oxycarbide (UCO) or UO2 fuel kernel, coated with buffer, inner pyrolytic carbon (IPyC), silicon carbide (SiC), and outer pyrolytic carbon (OPyC) layers. The SiC layer acts as the main containment barrier for the TRISO particle to retain the fission products, while the IPyC and OPyC layers provide additional barriers to the release of fission products, especially fission gases. During irradiation, phenomena like kernel swelling, buffer densification, and IPyC fracture may impact fuel performance. Post-irradiation microscopy on entire compact cross sections or samples of individual particles deconsolidated from compacts is often used to identify these irradiation-induced changes in morphology. However, each fuel compact generally contains thousands of TRISO particles. To get statistical information on these phenomena, it is cumbersome work if done manually. For example, to get information about swelling/densification behaviors of different layers or kernels after irradiation, researchers previously manually measured the perimeter of each TRISO layer in hundreds of particles after four rounds of iterative grinding and polishing encompassing more than 2000 cross-section images for a total of four fuel compacts. To attempt to reduce the subjectivity inherent in that process and accelerate data analysis, we conducted a study on the automatic TRISO layer segmentation on cross-sectional microscopic images using Convolutional Neural Networks (CNNs). CNNs are a class of machine learning algorithms specifically designed for processing structured grid data that have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we have generated the large irradiated TRISO layer dataset with more than 2000 cross-section TRISO microscopic images and the corresponding annotated images. Based on these annotated images, we have employed different CNNs for automatic segmentation of different TRISO layers. These include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net has the best performance in terms of intersection-over-union (IoU). Through the aid of these CNN models, we can expedite the analysis of TRISO particle cross-sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

Convolutional Neural Networks↗

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE↗

Estimating Fine-Resolution Shortwave Broadband Albedo of Croplands from Harmonized Landsat and Sentinel-2 Data

Altered surface albedo due to land-cover conversions and management is a significant driver of global climate change. Albedo can be directly measured at ground stations, and remote sensing data can be used to scale-up albedo values to regional and global levels. Some previous studies have retrieved fine-resolution (10–30 m) instantaneous albedo and coarse-resolution (500–1000 m) daily mean albedo from remote sensing data, but they all required the input of Moderate Resolution Imaging Spectroradiometer (MODIS) albedo information at 500-m resolution, and none have assembled both instantaneous and daily albedo based exclusively on fine-resolution satellite data. Here, to address this issue, we compiled 387 instantaneous and 346 daily albedo records using field net radiometer measurements from the bioenergy croplands at the W. K. Kellogg Biological Station in southwest Michigan. We then connected these albedo records with a suite of variables derived from harmonized Landsat and Sentinel-2 data through two machine learning algorithms (random forest regression and extreme gradient boosting) to retrieve clear-sky instantaneous and daily shortwave broadband albedo. The performance statistics indicate reasonable accuracy of model results [root-mean-square error (RMSE)] around or below 0.03 except for snow-covered surfaces), suggesting that the retrieval of both instantaneous and daily albedo based exclusively on fine-resolution satellite data is promising. To facilitate the use of fine-resolution albedo products at the global level, future efforts need to include more albedo records of diverse surface cover types, as well as to accurately model daily albedo for cloudy days to address the “clear-sky bias.”

Harmonized Landsat and Sentinel-2↗

Enabling scientific machine learning in MOOSE using Libtorch

A neural-network-based machine learning interface has been developed for the Multiphysics Object-Oriented Simulation Environment (MOOSE). The interface relies on Libtorch, the C++ front-end of PyTorch, and enables an online interaction between modern machine learning algorithms and all the existing simulation, modeling, and analysis processes available in MOOSE. New capabilities in MOOSE include the native generation and training of artificial neural networks together with options to load pretrained neural networks in TorchScript format. Furthermore, the MOOSE stochastic tools module (MOOSE-STM) has been enhanced with neural network-based surrogate and reduced-order model generation options for efficient stochastic analyses. Lastly, a reinforcement learning capability has been added to MOOSE-STM for the interactive control and optimization of complex multiphysics problems.

97 MATHEMATICS AND COMPUTING↗

Anomalously high elastic modulus of a poly(ethylene oxide)-based composite electrolyte

The practical use of lithium metal anodes in solid-state batteries requires a polymer membrane with high lithium-ion conductivity, thermal/electrochemical stability, and mechanical strength. The primary challenge is to effectively decouple the ionic conductivity and mechanical strength of the polymer electrolytes. We report a remarkably facile single step synthetic strategy based on in-situ crosslinking of poly(ethylene oxide) (xPEO) in the presence of a woven glass fiber (GF). Such a simple method yields composite polymer electrolytes (CPE) of anomalously high elastic modulus up to 2.5 GPa over a broad temperature range (20 °C – 245 °C) that has never been previously documented. An unsupervised machine learning algorithm, K-mean clustering analysis, was implemented on the hyperspectral Raman mapping at the xPEO/GF interface. Using such a unique means, we show for the first time that the promoted mechanical strength originates from xPEO and GF interactions through dynamic hydrogen and ionic bonding. High ionic conductivity is achieved by the addition plasticizer (e.g. tetraglyme), where trifluoromethanesulfonate anions are tethered to the xPEO matrix and Li + cations are favorably transported through coordination with the plasticizer. Further, stringent galvanostatic cycling tests indicates the CPE can be stably cycled for >3000 h in a Li-metal symmetric cell at a moderate temperature (nearly 1500 Coulombs/cm 2 Li equivalents), outperforming most of the PEO-based electrolytes. The GF reinforced CPE reported here has multifunctional uses, such as solid electrolytes for all solid-state batteries and membranes for redox-flow batteries. Although the focus of this study is on lithium-based batteries, the results are equally promising for other alkali metal based batteries such as sodium and potassium.

25 ENERGY STORAGE↗

Tailoring Molecular Space to Navigate Phase Complexity in Cs-Based Quasi-2D Perovskites via Gated-Gaussian-Driven High-Throughput Discovery

Cesium-based quasi-2D halide perovskites (HPs) offer promising functionalities and low-temperature manufacturability, suited to stable tandem photovoltaics. However, the chemical interplays between the molecular spacers and the inorganic building blocks during crystallization cause substantial phase complexities in the resulting matrices. To successfully optimize and implement the quasi-2D HP functionalities, a systematic understanding of spacer chemistry, along with the seamless navigation of the inherently discrete molecular space, is necessary. Herein, by utilizing high-throughput automated experimentation, the phase complexities in the molecular space of quasi-2D HPs are explored, thus identifying the chemical roles of the spacer cations on the synthesis and functionalities of the complex materials. Furthermore, a novel active machine learning algorithm leveraging a two-stage decision-making process, called gated Gaussian process Bayesian optimization is introduced, to navigate the discrete ternary chemical space defined with two distinctive spacer molecules. Through simultaneous optimization of photoluminescence intensity and stability that “tailors” the chemistry in the molecular space, a ternary-compositional quasi-2D HP film realizing excellent optoelectronic functionalities is demonstrated. Finally, this work not only provides a pathway for the rational and bespoke design of complex HP materials but also sets the stage for accelerated materials discovery in other multifunctional systems.

36 MATERIALS SCIENCE↗

Monte Carlo Dropout Uncertainty Quantification of Long Short-Term Memory Autoencoder Anomaly Detection in a Liquid Sodium Cold Trap

Advanced high-temperature fluid reactors, such as sodium-cooled fast reactors (SFRs) and molten salt–cooled reactors (MSCRs), require coolant purification systems to prevent fluid contamination and local freezing that can lead to plugging. Liquid sodium purification can be achieved with a cold trap, where the sodium temperature is reduced to a near-freezing point to precipitate out impurities. Automation of monitoring of the cold trap performance with machine learning algorithms can aid in early detection of incipient anomalies. An efficient approach to loss-of-coolant–type anomaly detection in a cold trap monitored with more than two dozen thermal-hydraulic sensors consists of a long short-term memory (LSTM) autoencoder. This work develops the uncertainty quantification of the LSTM autoencoder performance for cold trap anomaly detection using the Monte Carlo (MC) dropout method. The MC dropout methodology creates a distribution of sister distributions that all slightly differ from each other because of random neurons being turned off for testing. The variances of the sister network distributions are used to make an uncertainty interval. Our analysis shows that the uncertainty in the autoencoder performance is largest near the peak of the anomaly signal. Using the MC dropout method, we investigate the uncertainty in the anomaly detection with missing sensor inputs. This capability allows the reactor operator to evaluate resilience of the anomaly detection system and to make informed decisions about continuity of operation in the event of sensor failure.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Improving Subsurface Stress Characterization for Carbon Dioxide Storage Projects by Incorporating Machine Learning Techniques

The overall objective of this project is to develop a framework for reliable characterization and prediction of the state of stress in the overburden and underburden (including the basement) in CO 2 storage reservoirs using machine learning and integrated geomechanics and geophysical methods. Specifically, we propose to develop workflow encompassing of technologies and/or methods to predict stress and pressure changes due to CO 2 injection in an active tertiary recovery site and their impacts on subtle fault activation, fractures and occurrence of microseismic events and compare responses to field observations. In this project, we anticipate using dataset from the Farnsworth field Unit (FWU) which is operated by Purdure Petroleum. A novel elastic-waveform VSP inversion technique will be used to estimate high-resolution spatial and temporal changes of elastic moduli in CO 2 storage reservoirs, which will be combined with velocity-stress relationship derived from laboratory tests to obtain subsurface pressure and stress. Clustered microseismic data will be jointly inverted for improved focal mechanisms. Least-squares reverse-time migration of microseismic waveform data will be performed to directly image fracture/fault zones. Additionally, a deep neural network machine learning technique with convolutional and recurrent layers will be used for learning the spectro-temporal structures in microseismic waveforms. The results of this geotechnical data analysis will be integrated to develop a high-resolution 3D mechanical earth model extending from the overburden sealing formations to the underburden including the basement. Mechanical properties will be derived through integration of mechanical logs, tests, available results from chemo-mechanical laboratory tests, and elastic inversion of seismic data using a combination of Bayesian and stochastic methods as well as machine learning technique. Failure features (faults/fractures) will be represented and/or modeled based on seismic and core data analysis. A transient hydrodynamic-geomechanical model will be developed through coupling with the calibrated FWU reservoir simulation model. The full physics coupled model will be used to train a reduced order proxy model using machine learning algorithm for estimating stress which will then be used with appropriate constitutive relationships and forward seismological models to simulate pressure changes and induced microseismicity. An advanced optimization framework will be developed to perform a history match to minimize error between field observations and simulated. The history matched proxy model will be verified against the full-physics equivalent. The field observations that will be used in the coupled model calibration process include pressure/stress inverted from VSP, moment magnitude from microseismic analysis, real time downhole pressure measurements, production and injection data. Parameter sensitivity and uncertainty analysis will be performed to characterize the impact of model parameter uncertainty on stress estimates. The proposed project will have significant impact on future field implementation of the proposed technology. Because the project field site is an ongoing CO 2 EOR development, the value of the new technology will be demonstrated in an operational context and evaluated as a viable risk mitigation strategy. Cost/benefit will be evaluated together with the various commercial incentives for CO 2 sequestration available to oil and gas operators. The extensive available dataset and ongoing data acquisition under the SWP Phase III work plan provides flexibility for investigation of multiple approaches and reduces technical risk.

58 GEOSCIENCES↗

Ensemble‐Based, Large‐Eddy Reconstruction of Wind Turbine Inflow in a Near‐Stationary Atmospheric Boundary Layer Through Generative Artificial Intelligence

ABSTRACT To validate the second‐by‐second dynamics of turbines in field experiments, it is necessary to accurately reconstruct the winds going into the turbine. Current time‐resolved inflow reconstruction techniques estimate wind behavior in unobserved regions using relatively simple spectral‐based models of the atmosphere. Here, we develop a technique for time‐resolved inflow reconstruction that is rooted in a large‐eddy simulation model of the atmosphere. Our “large‐eddy reconstruction” technique blends observations and atmospheric model information through a diffusion model machine learning algorithm, allowing us to generate probabilistic ensembles of reconstructions for a single 10‐min observational period. Our generated inflows can be used directly by aeroelastic codes or as inflow boundary conditions in a large‐eddy simulation. We verify the second‐by‐second reconstruction capability of our technique in three synthetic field campaigns, finding positive Pearson correlation coefficient values () between ground‐truth and reconstructed streamwise velocity, as well as smaller positive correlation coefficient values for unobserved fields (spanwise velocity, vertical velocity, and temperature). We validate our technique in three real‐world case studies by driving large‐eddy simulations with reconstructed inflows and comparing to independent inflow measurements. The reconstructions are visually similar to measurements, follow desired power spectra properties, and track second‐by‐second behavior ().

17 WIND ENERGY↗

Communication-Avoiding and Memory-Constrained Sparse Matrix-Matrix Multiplication at Extreme Scale

Sparse matrix-matrix multiplication (SpGEMM) is a widely used kernel in various graph, scientific computing and machine learning algorithms. In this paper, we consider SpGEMMs performed on hundreds of thousands of processors generating trillions of nonzeros in the output matrix. Distributed SpGEMM at this extreme scale faces two key challenges: (1) high communication cost and (2) inadequate memory to generate the output. Furthermore, we address these challenges with an integrated communication-avoiding and memory-constrained SpGEMM algorithm that scales to 262,144 cores (more than 1 million hardware threads) and can multiply sparse matrices of any size as long as inputs and a fraction of output fit in the aggregated memory. As we go from 16,384 cores to 262,144 cores on a Cray XC40 supercomputer, the new SpGEMM algorithm runs 10x faster when multiplying large-scale protein-similarity matrices.

97 MATHEMATICS AND COMPUTING↗

Random Forests as a Viable Method to Select and Discover High-redshift Quasars

We present a method of selecting quasars up to redshift ≈6 with random forests, a supervised machine-learning method, applied to Pan-STARRS1 and WISE data. We find that, thanks to the increasing set of known quasars, we can assemble a training set that enables supervised machine-learning algorithms to become a competitive alternative to other methods up to this redshift. We present a candidate set for the redshift range 4.8–6.3, which includes the region around z = 5.5 where selecting quasars is difficult due to their photometric similarity to red and brown dwarfs. We demonstrate that, under our survey restrictions, we can reach a high completeness (66% ± 7% below redshift 5.6/83{sub -9}{sup +6}% above redshift 5.6) while maintaining a high selection efficiency (78{sub -8}{sup +10}%/94{sub -8}{sup +5}%). Our selection efficiency is estimated via a novel method based on the different distributions of quasars and contaminants on the sky. The final catalog of 515 candidates includes 225 known quasars. We predict the candidate catalog to contain additional 148{sub -33}{sup +41} new quasars below redshift 5.6 and 45{sub -8}{sup +5} above, and we make the catalog publicly available. Spectroscopic follow-up observations of 37 candidates led us to discover 20 new high redshift quasars (18 at 4.6 ≤ z ≤ 5.5, 2 z ~ 5.7). These observations are consistent with our predictions on efficiency. We argue that random forests can lead to higher completeness because our candidate set contains a number of objects that would be rejected by common color cuts, including one of the newly discovered redshift 5.7 quasars.

79 ASTRONOMY AND ASTROPHYSICS↗