Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

High-level hadronic tau lepton triggers of the CMS experiment in proton-proton collisions at √(s) = 13.6 TeV

The trigger system of the CMS detector is pivotal in the acquisition of data for physics measurements and searches. Studies of final states characterized by hadronic decays of tau leptons require the reconstruction and the identification of genuine tau leptons against quark- and gluon-initiated jets at the trigger level. This is a difficult task, particularly as improvements to the LHC have resulted in an increased number of interactions per bunch crossing in recent years. To address this challenge, a series of machine-learning algorithms with high identification efficiency and low computational cost have been incorporated into the high-level trigger for hadronically decaying tau leptons. In this paper, these developments and the trigger performance are summarized using data collected by the CMS experiment in proton-proton collisions at √(s) = 13.6 TeV in 2022–2023, corresponding to an integrated luminosity of 62 fb -1 .

Particle identification methods↗

Further adoption of conservation tillage can increase maize yields in the western US Corn Belt

Conservation tillage can reduce soil erosion, increase soil health, and decrease labor and fuel input costs. Despite these benefits, potential yield impacts remain an important concern for farmers considering adoption. Previous research suggests that conservation tillage is likely to have the largest yield benefits in more arid conditions, but a lack of field-level analyses across climatic, management and soil conditions limits confidence in such predictions. Satellite imagery provides the opportunity to monitor agricultural lands at sub-field resolution across large spatial scales and wide environmental gradients. Here we investigate the maize yield impacts of conservation tillage in the semi-arid western US Corn Belt, using sub-field resolution datasets on tillage practices and crop yields derived from satellite data spanning four states (Nebraska, Kansas, South Dakota, and North Dakota) between 2008 and 2020. On these datasets, we estimate heterogenous yield outcomes for several thousand maize fields across gradients in climate, soil quality and irrigation status by using a causal forests analysis, an adaptation of the random forests machine-learning algorithm for causal inference on observational data. We find that long-term adoption of conservation tillage increased rainfed maize yields by an average of 9.9% in the region. Impacts on irrigated yields were small and not statistically significant. These results, along with an analysis of variables related to greater than average yield benefits, indicate that improved water infiltration and retention are the primary reasons for conservation tillage benefits. Despite yield benefits, many fields estimated to see increased yields under long term low till have not adopted the practice. Therefore, we identify specific counties likely to benefit most from increased levels of adoption. Our results strengthen the understanding of the impacts of conservation agriculture on crop yields and help define environments and counties most likely to benefit from conservation tillage.

54 ENVIRONMENTAL SCIENCES↗

High-dimensional multi-fidelity Bayesian optimization for quantum control

Abstract We present the first multi-fidelity Bayesian optimization (BO) approach for solving inverse problems in the quantum control of prototypical quantum systems. Our approach automatically constructs time-dependent control fields that enable transitions between initial and desired final quantum states. Most importantly, our BO approach gives impressive performance in constructing time-dependent control fields, even for cases that are difficult to converge with existing gradient-based approaches. We provide detailed descriptions of our machine learning methods as well as performance metrics for a variety of machine learning algorithms. Taken together, our results demonstrate that BO is a promising approach to efficiently and autonomously design control fields in general quantum dynamical systems.

Computer Science↗

Journey over Destination: Dynamic Sensor Placement Enhances Generalization

Reconstructing complex, high-dimensional global fields from limited data points is a challenge across various scientific and industrial domains. This is particularly important for recovering spatio-temporal fields using sensor data from, for example, laboratory-based scientific experiments, weather forecasting, or drone surveys. Given the prohibitive costs of specialized sensors and the inaccessibility of
certain regions of the domain, achieving full field coverage is typically not feasible. Therefore, the development of machine learning algorithms trained to reconstruct fields given a limited dataset is of critical importance. In this study, we introduce a general
approach that employs moving sensors to enhance data exploitation during the training of an attention based neural network, thereby improving field reconstruction. The training of sensor locations is accomplished using an end-to-end workflow, ensuring
differentiability in the interpolation of field values associated to the sensors, and is simple to implement using differentiable programming. Additionally, we have incorporated a correction mechanism to prevent sensors from entering invalid regions within the domain. We evaluated our method using two distinct datasets; the results show that our approach enhances learning, as evidenced by improved test scores.

54 ENVIRONMENTAL SCIENCES↗

Deep probabilistic direction prediction in 3D with applications to directional dark matter detectors

Abstract We present the first method to probabilistically predict 3D direction in a deep neural network model. The probabilistic predictions are modeled as a heteroscedastic von Mises-Fisher distribution on the sphere S 2 , giving a simple way to quantify aleatoric uncertainty. This approach generalizes the cosine distance loss which is a special case of our loss function when the uncertainty is assumed to be uniform across samples. We develop approximations required to make the likelihood function and gradient calculations stable. The method is applied to the task of predicting the 3D directions of electrons, the most complex signal in a class of experimental particle physics detectors designed to demonstrate the particle nature of dark matter and study solar neutrinos. Using simulated Monte Carlo data, the initial direction of recoiling electrons is inferred from their tortuous trajectories, as captured by the 3D detectors. For 40 keV electrons in a 70% He 30% CO 2 gas mixture at STP, the new approach achieves a mean cosine distance of 0.104 (26 ∘ ) compared to 0.556 (64 ∘ ) achieved by a non-machine learning algorithm. We show that the model is well-calibrated and accuracy can be increased further by removing samples with high predicted uncertainty. This advancement in probabilistic 3D directional learning could increase the sensitivity of directional dark matter detectors.

Computer Science↗

Global Sensitivity Analysis of a Reactive Transport Model for Mineral Scale Formation During Hydraulic Fracturing

Injection of water-based hydraulic fracturing fluid (HFF) into tight shale gas/oil formations can increase formation permeability and enhance production rates, but this process frequently causes mineral scale formation that can occlude pore space and hinder flow. To identify the most important factors that control the formation of mineral scales, we applied a novel global sensitivity analysis method—distance-based generalized sensitivity analysis (DGSA)—to a reactive transport model (RTM) that was previously built and calibrated to simulate precipitation of barite [BaSO4] and iron (hydr)oxide [Fe(OH) 3 ] in shale matrices and on fracture surfaces. Reactive transport simulations were run with model parameters randomly sampled based on assigned uncertainties. Modeling results for barite and Fe(OH)3 formation were clustered using machine-learning algorithms. A list of ranked critical input parameters was obtained after statistical quantification of cumulative distribution functions of input parameters. We found that barite formation is most sensitive to the rate of sulfate ion generation, which is determined by the pyrite dissolution rate coefficient and oxidant availability. In addition, barite formation is sensitive to the initial amounts of barite in HFF and shale, followed by barite thermodynamics/kinetics. For Fe(OH) 3 formation, the ranked factors are Fe(OH)3 precipitation rate coefficients, initial HFF pH, initial Fe(OH) 3 amount in HFF, and oxidant availability. Overall, our results provide insights into managing mineral scale formation during hydraulic fracturing to enhance production. Meanwhile, this study serves as an example of global sensitivity analysis of RTMs using the efficient, straightforward, and open-source DGSA method.

58 GEOSCIENCES↗

EcoPLOT: dynamic analysis of biogeochemical data

Motivation: We have created EcoPLOT (parameterized linkage of omics-driven technologies), a web-app for the dynamic, interactive analysis of biogeochemical datasets that combines state-of-the-art analysis tools to statistically and graphically explore environmental, geochemical and microbiome datasets. Using the iterative random forest, a machine learning algorithm, EcoPLOT allows for the de novo discovery of drivers which exhibit significant impact on plant, microbial or soil dynamics. Availability and implementation: EcoPLOT is built entirely within the R language. It can be accessed through any system where R is installed, including Windows, Mac and most Linux systems. EcoPLOT is free to use and can be accessed at https://github.com/cdsanchez18/EcoPLOT.

59 BASIC BIOLOGICAL SCIENCES↗

Transforming microseismic clouds into near real-time visualization of the growing hydraulic fracture

SUMMARY Microseismic observations during unconventional reservoir stimulation are typically seen as a proxy for clusters of hydraulic fractures and the extent of the stimulated reservoir. Such straightforward interpretation is often misleading and fails to provide a physically reasonable image of the fracturing process. This paper demonstrates the application of a physics-based machine learning algorithm which enables a rapid and accurate fracture mapping from the microseismic data. Our training and validation data set relies on a history-matched geomechanical modelling workflow implemented in GEOS software for the Hydraulic Fracturing Test Site 1 (HFTS-1) project. For this study we augmented the simulated fracture growth through geostatistical modelling of induced seismicity, so that the synthetic microseismic catalogue matches the main statistical properties of the field observations. We formulated the problem of mapping the actual fracture in the clutter of events to parallel common video segmentation workflows: several past video frames (microseismic density snapshots) are passed through a deep convolutional network to classify whether a given voxel is associated with a fracture or intact rock. We found that for accurate fracture mapping, the network’s input and architecture must be augmented to incorporate the fluid injection parameters (pressure, rate, concentration of proppant, and location of the perforation within the cluster). The error rate for the network reached as little as 10 per cent of the fracture area, while a conventional microseismic interpretation approach yielded ∼300 per cent. Our approach also yields must faster predictions than conventional methods (minutes instead of weeks), and could enable engineers to make rapid decisions regarding engineering parameters (pumping rate, viscosity) in real time during stimulation.

58 GEOSCIENCES↗

Dark Energy Survey Year 3 results: galaxy sample for BAO measurement

Here, we present and validate the galaxy sample used for the analysis of the baryon acoustic oscillation (BAO) signal in the Dark Energy Survey (DES) Y3 data. The definition is based on a colour and redshift-dependent magnitude cut optimized to select galaxies at redshifts higher than 0.5, while ensuring a high-quality determination. The sample covers ~ 4100 square degrees to a depth of i = 22.3 (AB) at 10σ. It contains 7,031,993 galaxies in the redshift range from z = 0.6 to 1.1, with a mean effective redshift of 0.835. Redshifts are estimated with the machine learning algorithm DNF, and are validated using the VIPERS PDR2 sample. We find a mean redshift bias of z bias ~ 0.01 and a mean uncertainty, in units of 1 + z, of σ 68 ~ 0.03$. We evaluate the galaxy population of the sample, showing it is mostly built upon Elliptical to Sbc types. Furthermore, we find a low level of stellar contamination of ≲ 4 %. We present the method used to mitigate the effect of spurious clustering coming from observing conditions and other large-scale systematics. We apply it to the BAO sample and calculate weights that are used to get a robust estimate of the galaxy clustering signal. This paper is one of a series dedicated to the analysis of the BAO signal in DES Y3. In the companion papers, we present the galaxy mock catalogues used to calibrate the analysis and the angular diameter distance constraints obtained through the fitting to the BAO scale.

79 ASTRONOMY AND ASTROPHYSICS↗

Angular clustering properties of the DESI QSO target selection using DR9 Legacy Imaging Surveys

ABSTRACT The quasar target selection for the upcoming survey of the Dark Energy Spectroscopic Instrument (DESI) will be fixed for the next 5 yr. The aim of this work is to validate the quasar selection by studying the impact of imaging systematics as well as stellar and galactic contaminants, and to develop a procedure to mitigate them. Density fluctuations of quasar targets are found to be related to photometric properties such as seeing and depth of the Data Release 9 of the DESI Legacy Imaging Surveys. To model this complex relation, we explore machine learning algorithms (random forest and multilayer perceptron) as an alternative to the standard linear regression. Splitting the footprint of the Legacy Imaging Surveys into three regions according to photometric properties, we perform an independent analysis in each region, validating our method using extended Baryon Oscillation Spectroscopic Survey (eBOSS) EZ-mocks. The mitigation procedure is tested by comparing the angular correlation of the corrected target selection on each photometric region to the angular correlation function obtained using quasars from the Sloan Digital Sky Survey (SDSS) Data Release 16. With our procedure, we recover a similar level of correlation between DESI quasar targets and SDSS quasars in two-thirds of the total footprint and we show that the excess of correlation in the remaining area is due to a stellar contamination that should be removed with DESI spectroscopic data. We derive the Limber parameters in our three imaging regions and compare them to previous measurements from SDSS and the 2dF QSO Redshift Survey.

79 ASTRONOMY AND ASTROPHYSICS↗

via machinae : Searching for stellar streams using unsupervised machine learning

ABSTRACT We develop a new machine learning algorithm, via machinae, to identify cold stellar streams in data from the Gaia telescope. via machinae is based on ANODE, a general method that uses conditional density estimation and sideband interpolation to detect local overdensities in the data in a model agnostic way. By applying ANODE to the positions, proper motions, and photometry of stars observed by Gaia, via machinae obtains a collection of those stars deemed most likely to belong to a stellar stream. We further apply an automated line-finding method based on the Hough transform to search for line-like features in patches of the sky. In this paper, we describe the via machinae algorithm in detail and demonstrate our approach on the prominent stream GD-1. Though some parts of the algorithm are tuned to increase sensitivity to cold streams, the via machinae technique itself does not rely on astrophysical assumptions, such as the potential of the Milky Way or stellar isochrones. This flexibility suggests that it may have further applications in identifying other anomalous structures within the Gaia data set, for example debris flow and globular clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

AutoEnRichness: A hybrid empirical and analytical approach for estimating the richness of galaxy clusters

ABSTRACT We introduce AutoEnRichness, a hybrid approach that combines empirical and analytical strategies to determine the richness of galaxy clusters (in the redshift range of 0.1 ≤ z ≤ 0.35) using photometry data from the Sloan Digital Sky Survey Data Release 16, where cluster richness can be used as a proxy for cluster mass. In order to reliably estimate cluster richness, it is vital that the background subtraction is as accurate as possible when distinguishing cluster and field galaxies to mitigate severe contamination. AutoEnRichness is comprised of a multistage machine learning algorithm that performs background subtraction of interloping field galaxies along the cluster line of sight and a conventional luminosity distribution fitting approach that estimates cluster richness based only on the number of galaxies within a magnitude range and search area. In this proof-of-concept study, we obtain a balanced accuracy of 83.20 per cent when distinguishing between cluster and field galaxies as well as a median absolute percentage error of 33.50 per cent between our estimated cluster richnesses and known cluster richnesses within r200. In the future, we aim for AutoEnRichness to be applied on upcoming large-scale optical surveys, such as the Legacy Survey of Space and Time and Euclid, to estimate the richness of a large sample of galaxy groups and clusters from across the halo mass function. This would advance our overall understanding of galaxy evolution within overdense environments as well as enable cosmological parameters to be further constrained.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep learning methods for obtaining photometric redshift estimations from images

ABSTRACT Knowing the redshift of galaxies is one of the first requirements of many cosmological experiments, and as it is impossible to perform spectroscopy for every galaxy being observed, photometric redshift (photo-z) estimations are still of particular interest. Here, we investigate different deep learning methods for obtaining photo-z estimates directly from images, comparing these with ‘traditional’ machine learning algorithms which make use of magnitudes retrieved through photometry. As well as testing a convolutional neural network (CNN) and inception-module CNN, we introduce a novel mixed-input model that allows for both images and magnitude data to be used in the same model as a way of further improving the estimated redshifts. We also perform benchmarking as a way of demonstrating the performance and scalability of the different algorithms. The data used in the study comes entirely from the Sloan Digital Sky Survey (SDSS) from which 1 million galaxies were used, each having 5-filtre (ugriz) images with complete photometry and a spectroscopic redshift which was taken as the ground truth. The mixed-input inception CNN achieved a mean squared error (MSE) =0.009, which was a significant improvement ($30{{\ \rm per\ cent}}$) over the traditional random forest (RF), and the model performed even better at lower redshifts achieving a MSE = 0.0007 (a $50{{\ \rm per\ cent}}$ improvement over the RF) in the range of z < 0.3. This method could be hugely beneficial to upcoming surveys, such as Euclid and the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), which will require vast numbers of photo-z estimates produced as quickly and accurately as possible.

79 ASTRONOMY AND ASTROPHYSICS↗

The Human Phenotype Ontology in 2024: phenotypes around the world

The Human Phenotype Ontology (HPO) is a widely used resource that comprehensively organizes and defines the phenotypic features of human disease, enabling computational inference and supporting genomic and phenotypic analyses through semantic similarity and machine learning algorithms. The HPO has widespread applications in clinical diagnostics and translational research, including genomic diagnostics, gene-disease discovery, and cohort analytics. In recent years, groups around the world have developed translations of the HPO from English to other languages, and the HPO browser has been internationalized, allowing users to view HPO term labels and in many cases synonyms and definitions in ten languages in addition to English. Since our last report, a total of 2239 new HPO terms and 49235 new HPO annotations were developed, many in collaboration with external groups in the fields of psychiatry, arthrogryposis, immunology and cardiology. The Medical Action Ontology (MAxO) is a new effort to model treatments and other measures taken for clinical management. Finally, the HPO consortium is contributing to efforts to integrate the HPO and the GA4GH Phenopacket Schema into electronic health records (EHRs) with the goal of more standardized and computable integration of rare disease data in EHRs.

59 BASIC BIOLOGICAL SCIENCES↗

Retrieval of full angular- and energy-dependent complex transition dipoles in the molecular frame from laser-induced high-order harmonic signals with aligned molecules

High-order harmonic signals generated in molecules are the consequence of coherent summation of complex laser-induced transition dipoles $\textit{d}(θ, ω)$ with each fixed-in-space molecule; here $\theta$ is the angle of the molecular axis with respect to the laser polarization axis and $\omega$ is the harmonic energy. In the so-called rotational coherent spectroscopy, it is proposed to extract the fixed-in-space $\textit{d}(θ, ω)$ in the molecular frame by measuring harmonics generated by a probing laser from the rotational molecular wave packets that have been prepared by a prior pump laser. By varying the time delay between the two lasers, methods have been utilized to extract the $\theta$ dependence of both the amplitude and phase of each individual harmonic, but the relative phase between harmonics cannot be retrieved. Here we report that this limitation can be removed. It requires the additional measurement of harmonic spectra versus the pump-probe angles at one fixed time delay. The two-dimensional input harmonic data (time-delay and pump-probe angle) are then used to retrieve the full complex transition dipole $\textit{d}(θ, ω)$ using a retrieval method based on machine learning algorithms. Finally, we demonstrate this method on N 2 and CO 2 molecules.

74 ATOMIC AND MOLECULAR PHYSICS↗

Machine-learning prediction for quasiparton distribution function matrix elements

There have been rapid developments in the direct calculation in lattice QCD (LQCD) of the Bjorken-x dependence of hadron structure through large-momentum effective theory (LaMET). LaMET overcomes the previous limitation of LQCD to moments (that is, integrals over Bjorken x) of hadron structure, allowing LQCD to directly provide the kinematic regions where the experimental values are least known. LaMET requires large-momentum hadron states to minimize its systematics and allow us to reach small- x reliably. This means that very fine lattice spacing to minimize lattice artifacts at order $(P_za)^n$ will become crucial for next-generation LaMET-like structure calculations. Furthermore, such calculations require operators with long Wilson-link displacements, especially in finer lattice units, increasing the communication costs relative to that of the propagator inversion. In this work, we explore whether machine-learning algorithms can make predictions of correlators to reduce the computational cost of these LQCD calculations. We consider two algorithms, gradient-boosting decision tree and linear models, applied to LaMET data, the matrix elements needed to determine the kaon and $η_s$ unpolarized parton distribution functions (PDFs), meson distribution amplitude (DA), and the nucleon gluon PDF. We find that both algorithms can reliably predict the target observables with different prediction accuracy and systematic errors. The predictions from smaller displacement z to larger ones work better than those for momentum p due to the higher correlation among the data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Search for long-lived particles using displaced vertices and missing transverse momentum in proton-proton collisions at s = 13 TeV

A search for the production of long-lived particles in proton-proton collisions at a center-of-mass energy of 13 TeV at the CERN LHC is presented. The search is based on data collected by the CMS experiment in 2016–2018, corresponding to a total integrated luminosity of 137 fb − 1 . This search is designed to be sensitive to long-lived particles with mean proper decay lengths between 0.1 and 1000 mm, whose decay products produce a final state with at least one displaced vertex and missing transverse momentum. A machine learning algorithm, which improves the background rejection power by more than an order of magnitude, is applied to improve the sensitivity. The observation is consistent with the standard model background prediction, and the results are used to constrain split supersymmetry (SUSY) and gauge-mediated SUSY breaking models with different gluino mean proper decay lengths and masses. This search is the first CMS search that shows sensitivity to hadronically decaying long-lived particles from signals with mass differences between the gluino and neutralino below 100 GeV. It sets the most stringent limits to date for split-SUSY models and gauge-mediated SUSY breaking models with gluino proper decay length less than 6 mm. © 2024 CERN, for the CMS Collaboration 2024 CERN

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Characterization of the astrophysical diffuse neutrino flux using starting track events in IceCube

In this article, a measurement of the diffuse astrophysical neutrino spectrum is presented using IceCube data collected from 2011-2022 (10.3 years). We developed novel detection techniques to search for events with a contained vertex and exiting track induced by muon neutrinos undergoing a charged-current interaction. Searching for these starting track events allows us to not only more effectively reject atmospheric muons but also atmospheric neutrino backgrounds in the southern sky, opening a new window to the sub-100 TeV astrophysical neutrino sky. The event selection is constructed using a dynamic starting track veto and machine learning algorithms. We use this data to measure the astrophysical diffuse flux as a single power law flux (SPL) with a best-fit spectral index of γ=2.58$_{-0.09}^{+0.10}$ and per-flavor normalization of $\phi$$_{per-flavor}^{Astro}$=1.68$_{-0.22}^{+0.19}$×10 -18 ×GeV -1 cm -2 s -1 sr -1 (at 100 TeV). The sensitive energy range for this dataset is 3-550 TeV under the SPL assumption. This data was also used to measure the flux under a broken power law, however we did not find any evidence of a low energy cutoff.

79 ASTRONOMY AND ASTROPHYSICS↗