Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “massive datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Sensitivity to low-mass WIMPs with an improved liquid argon ionization response model within the DarkSide program

Dark matter detection experiments using liquid argon rely on a precise characterization of the ionization response to nuclear recoils, especially in the keV energy range relevant for light dark matter interactions. In this work, we present a comprehensive analysis that combines new measurements from the ReD setup, part of the DarkSide experimental program, with calibration data from DarkSide-50, as well as results from the ARIS and SCENE experiments. These combined datasets enable improved constraints on atomic screening effects in the modeling of the ionization response of liquid argon to nuclear recoils. The analysis is performed within the Thomas-Imel recombination framework adopted in previous DarkSide studies, and is here further constrained by the inclusion of ReD data, which allow the screening function to be determined from calibration measurements. By including the updated ionization model into the DarkSide-50 analysis framework, we obtain stronger exclusion limits on low-mass weakly interacting massive particle (WIMP) interactions, setting new world-leading constraints in the 1 – 3 GeV / c 2 WIMP mass range. Finally, we recast the sensitivity projections for the upcoming DarkSide-20k detector, demonstrating a significantly enhanced discovery potential for low-mass dark matter candidates.

Acerbi, F. [Fond. Bruno Kessler, Trento]↗

X-Ray Gas Temperatures in the Arc Clusters MS0440+204 and MS0302+1658

The cluster of galaxies MS0440+02, originally discovered through its X-ray emission, was part of an optical observational program to search for arcs and arclets in a complete sample of X-ray luminous, medium-distant clusters of galaxies. Mauna Kea CCD images of MS0440+02 showed a remarkable optical morphology. The core of the cluster contains 6 bright galaxies and numerous fainter ones embedded in a low surface brightness halo. Besides, MS0440+02 is the most spectacular example that we have found of an arc system in a compact condensed cluster, with arcs symmetrically distributed to draw almost perfect circles around the cluster center. Giant arcs are magnified images of distant galaxies, gravitationally distorted by massive foreground clusters. It is of great importance to compare the results of the lensing studies with those derived from X-ray observations, as the two are independent methods of studying the mass distribution. Thus MS0440+02 was the ideal target to obtain temperature measurement with ASCA and good spatial resolution X-ray observations with ROSAT. The X-ray data have been used in conjunction with Hubble Space Telescope observations to put more stringent constrains on the mass estimates. Most of the different wavelength datasets have been reduced and analyzed. Mass determinations have been separately obtained from galaxy virial motions and X-ray profile fitting using the cluster gas temperature as measured by the ASCA satellite. Assuming that the hot gas is in hydrostatic equilibrium and in a spherical potential, we find from the X-ray data a mass distribution profile that is well described by a Beta model. From the multiple images formed by gravitational lensing (HST data) using the modelling of the gravitational lensed arcs, we have derived Beta model. To reconcile the mass estimates we have explored the possibility of having a supercluster surrounding the MOS0440 cluster, that is a model with two isothermal spheres, one embedded inside the other. These results have been published or are in press.

Gioia, Isabella M.↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Dark Matter Search Results from 4.2 Tonne−Years of Exposure of the LUX-ZEPLIN (LZ) Experiment

We report results of a search for nuclear recoils induced by weakly interacting massive particle (WIMP) dark matter using the LUX-ZEPLIN (LZ) two-phase xenon time projection chamber. This analysis uses a total exposure of 4.2 ±0.1 tonne-years from 280 live days of LZ operation, of which 3.3 ± 0.1 tonne-years and 220 live days are new. A technique to actively tag background electronic recoils from 214 Pb 𝛽 decays is featured for the first time. Enhanced electron-ion recombination is observed in two-neutrino double electron capture decays of 124 Xe, representing a noteworthy new background. After removal of artificial signal-like events injected into the dataset to mitigate analyzer bias, we find no evidence for an excess over expected backgrounds. World-leading constraints are placed on spin-independent (SI) and spin-dependent WIMP-nucleon cross sections for masses ≥9 GeV/𝑐 2 . The strongest SI exclusion set is 2.2×10 −48 cm 2 at the 90% confidence level and the best SI median sensitivity achieved is 5.1 ×10 −48 cm 2 , both for a mass of 40 GeV/𝑐 2 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Coupling Between Tropical Meteorology, Aerosol Lifecycle, Convection, and Radiation, During the Clouds, Aerosol and Monsoon Processes Philippines Experiment (CAMP2Ex)

The NASA Cloud, Aerosol, and Monsoon Processes Philippines Experiment (CAMP2Ex) employed the NASA P-3, Stratton Park Engineering Company (SPEC) Learjet 35, and a host of satellites and surface sensors to characterize the coupling of aerosol processes, cloud physics, and atmospheric radiation within the Maritime Continent’s complex southwest monsoonal environment. Conducted in the late summer of 2019 from Luzon Philippines in conjunction with the Office of Naval Research Propagation of Intraseasonal Tropical OscillatioNs (PISTON) experiment with its R/VSally Ride stationed in the North Western Tropical Pacific, CAMP2Ex documented diverse biomass burning, industrial and natural aerosol populations and their interactions with small to congestus convection. The 2019 season exhibited El Nino and associated drought, high biomass burning emissions, and an early monsoon transition allowing for observation of pristine to massively polluted environments as they advected through intricate diurnal mesoscale and radiative environments into the monsoonal trough. CAMP2Ex’s preliminary results indicate 1) increasing aerosol loadings tend to invigorate congestus convection in height and increase liquid water paths; 2) lidar, polarimetry, and geostationary Advanced Himawari Imager remote sensing sensors have skill in quantifying diverse aerosol and cloud properties and their interaction; and 3) high resolution remote sensing technologies are able to greatly improve our ability to evaluate the radiation budget in complex cloud systems. Through the development of innovative informatics technologies, CAMP2Ex provides a benchmark dataset of an environment of extremes for the study of aerosol, cloud and radiation processes as well as a crucible for the design of future observing systems.

CAMP2Ex↗

NeMO-Net: The Neural Multi-Modal Observation and Training Network for Global Coral Reef Assessment

In the past decade, coral reefs worldwide have experienced unprecedented stresses due to climate change, ocean acidification, and anthropomorphic pressures, instigating massive bleaching and die-off of these fragile and diverse ecosystems. Furthermore, remote sensing of these shallow marine habitats is hindered by ocean wave distortion, refraction and optical attenuation, leading invariably to data products that are often of low resolution and signal-to-noise (SNR) ratio. However, recent advances in UAV and Fluid Lensing technology have allowed us to capture multispectral 3D imagery of these systems at sub-cm scales from above the water surface, giving us an unprecedented view of their growth and decay. Exploiting the fine-scaled features of these datasets, machine learning methods such as MAP, PCA, and SVM can not only accurately classify the living cover and morphology of these reef systems (below 8 percent error), but are also able to map the spectral space between airborne and satellite imagery, augmenting and improving the classification accuracy of previously low-resolution datasets. We are currently implementing NeMO-Net, the first open-source deep convolutional neural network (CNN) and interactive active learning and training software to accurately assess the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology. NeMO-Net will be built upon the QGIS platform to ingest UAV, airborne and satellite datasets from various sources and sensor capabilities, and through data-fusion determine the coral reef ecosystem makeup globally at unprecedented spatial and temporal scales. To achieve this, we will exploit virtual data augmentation, the use of semi-supervised learning, and active learning through a tablet platform allowing for users to manually train uncertain or difficult to classify datasets. The project will make use of Pythons extensive libraries for machine learning, as well as extending integration to GPU and High-End Computing Capability (HECC) on the Pleiades supercomputing cluster, located at NASA Ames. The project is being supported by NASAs Earth Science Technology Office (ESTO) Advanced Information Systems Technology (AIST-16) Program.

NeMO-Net↗

NeMO-Net The Neural Multi-Modal Observation Training Network for Global Coral Reef Assessment

In the past decade, coral reefs worldwide have experienced unprecedented stresses due to climate change, ocean acidification, and anthropomorphic pressures, instigating massive bleaching and die-off of these fragile and diverse ecosystems. Furthermore, remote sensing of these shallow marine habitats is hindered by ocean wave distortion, refraction and optical attenuation, leading invariably to data products that are often of low resolution and signal-to-noise (SNR) ratio. However, recent advances in UAV and Fluid Lensing technology have allowed us to capture multispectral 3D imagery of these systems at sub-cm scales from above the water surface, giving us an unprecedented view of their growth and decay. Exploiting the fine-scaled features of these datasets, machine learning methods such as MAP, PCA, and SVM can not only accurately classify the living cover and morphology of these reef systems (below 8 error), but are also able to map the spectral space between airborne and satellite imagery, augmenting and improving the classification accuracy of previously low-resolution datasets.We are currently implementing NeMO-Net, the first open-source deep convolutional neural network (CNN) and interactive active learning and training software to accurately assess the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology. NeMO-Net will be built upon the QGIS platform to ingest UAV, airborne and satellite datasets from various sources and sensor capabilities, and through data-fusion determine the coral reef ecosystem makeup globally at unprecedented spatial and temporal scales. To achieve this, we will exploit virtual data augmentation, the use of semi-supervised learning, and active learning through a tablet platform allowing for users to manually train uncertain or difficult to classify datasets. The project will make use of Pythons extensive libraries for machine learning, as well as extending integration to GPU and High-End Computing Capability (HECC) on the Pleiades supercomputing cluster, located at NASA Ames. The project is being supported by NASAs Earth Science Technology Office (ESTO) Advanced Information Systems Technology (AIST-16) Program.

Remote Sensin↗

Investigating performance and variability of NIF ICF experiments with deep learning

The parameter space involved in designing an inertial confinement fusion shot at the National Ignition Facility (NIF) is massively multi-dimensional and the cost of a single shot makes a comprehensive set of sensitivity studies in the laboratory impractical. The use of machine learning to overcome these challenges has gained popularity and has had several successful applications by the scientific community. We extend on these efforts by training a neural network (NN) on information about the experimental design, engineering elements, and drive asymmetry to predict with uncertainty the neutron yield of an experiment. We find the measured and model predicted values are in good agreement, with an R 2 value of 0.91 for a randomly selected test dataset. Almost all the predicted 95% credible intervals contain the corresponding measured value for both training and test datasets. We identify correlations picked up by the NN between the shot design, yield, and variability and use them to motivate shot sensitivity studies. The first shot to exceed the Lawson-like ignition criteria (N210808) was conducted at the NIF and subsequent shots studied the design’s robustness. In a follow-up shot to N210808, our model predicts capsule quality to be the main performance degradation mechanism that prevented the shot from repeating previous performance levels. Shot N221204 was the first shot to exceed a target energy gain of 1. Our model predicts increased yield with reduced coast time for a N221204 study and greater variability for designs with lower peak powers at constant yield. The model’s fast prediction speed and uncertainty prediction are useful for identifying interesting design paths that could warrant further investigation with conventional simulations to search for robust high yield designs.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NeMO-Net - The Neural Multi-Modal Observation & Training Network for Global Coral Reef Assessment

In the past decade, coral reefs worldwide have experienced unprecedented stresses due to climate change, ocean acidification, and anthropomorphic pressures, instigating massive bleaching and die-off of these fragile and diverse ecosystems. Furthermore, remote sensing of these shallow marine habitats is hindered by ocean wave distortion, refraction and optical attenuation, leading invariably to data products that are often of low resolution and signal-to-noise (SNR) ratio. However, recent advances in UAV and Fluid Lensing technology have allowed us to capture multispectral 3D imagery of these systems at sub-cm scales from above the water surface, giving us an unprecedented view of their growth and decay. By combining spatial and spectral information from varying resolutions, we seek to augment and improve the classification accuracy of previously low-resolution datasets at large temporal scales.NeMO-Net, the first open-source deep convolutional neural network (CNN) and interactive learning and training software, currently being developed at NASA Ames, is aimed at assessing the present and past dynamics of coral reef ecosystems through determination of percent living cover and morphology. The latest iteration uses fully convolutional networks to segment and identify coral imagery taken by UAVs and satellites, including WorldView-2 and Sentinel. We present results taken from the Indian Ocean where classification accuracy has exceeded 91% for 24 geomorphological classes given ample training data. In addition, we utilize deep Laplacian Pyramid Super-Resolution Networks (LapSRN) to reconstruct high resolution information from low resolution imagery, trained from various UAV and satellite datasets. Finally, in the case of insufficient training data, we have developed an interactive online platform that allows users to easily segment and submit their classifications, which has been integrated with the current NeMO-Net workflow. Specifically, we present results from the Fiji islands in which preliminary user data has allowed for the accurate identification of 9 separate classes, despite issues such as cloud shadowing and spectral variation. The project is being supported by NASA's Earth Science Technology Office (ESTO) Advanced Information Systems Technology (AIST-16) Program.

Neural↗

A Web Architecture to Geographically Interrogate CHIRPS Rainfall and eMODIS NDVI for Land Use Change

Monitoring of rainfall and vegetation over the continent of Africa is important for assessing the status of crop health and agriculture, along with long‐term changes in land use change. These issues can be addressed through examination of long‐term precipitation (rainfall) data sets and remote sensing of land surface vegetation and land use types. Two products have been used previously to address these goals: the Climate Hazard Group Infrared Precipitation with Stations (CHIRPS) rainfall data, and multi‐day composites of Normalized Difference Vegetation Index (NDVI) from the USGS eMODIS product. Combined, these are very large data sets that require unique tools and architecture to facilitate a variety of data analysis methods or data exploration by the end user community. To address these needs, a web‐enabled system has been developed to allow end‐users to interrogate CHIRPS rainfall and eMODIS NDVI data over the continent of Africa. The architecture allows end‐users to use custom defined geometries, or the use of predefined political boundaries in their interrogation of the data. The massive amount of data interrogated by the system allows the end‐users with only a web browser to extract vital information in order to investigate land use change and its causes. The system can be used to generate daily, monthly and yearly averages over a geographical area and range of dates of interest to the user. It also provides analysis of trends in precipitation or vegetation change for times of interest. The data provided back to the end‐user is displayed in graphical form and can be exported for use in other, external tools. The development of this tool has significantly decreased the investment and requirements for end‐users to use these two important datasets, while also allowing the flexibility to the end‐user to limit the search to the area of interest.

Burks, Jason E.↗

The Future of NASA Earth Science in the Commercial Cloud: Challenges and Opportunities

NASA produces a large volume and variety of data products that are used every day to support research, decision making, and education. The widespread use of NASA’s Earth Science data is enabled by NASA’s Earth Science Data System (ESDS) program, which oversees the archiving and distribution of these data and invests in the development of new data systems and tools. However, NASA’s current approach to Earth Science data distribution — based on distributed institutional archives with individual on-premises high-performance computing capabilities — faces some significant challenges, including massive increases in data volume from upcoming missions, a greater need for transdisciplinary science that synthesizes many different kinds of observations, and a push to make science more open, inclusive, and accessible. To address these challenges, NASA is aggressively migrating its Earth Science data and related tools and services into the commercial cloud. Migration of data into the commercial cloud can significantly improve NASA’s existing data system capabilities by (1) providing more flexible options for storage and compute (including rapid, as-needed access to state-of-the-art capabilities); (2) by centralizing and standardizing data access, which gives all of NASA’s institutional data centers access to all of each other’s datasets; and (3) by facilitating “analysis-in-place”, whereby users can bring their own computational workflows and tools to the data rather than having to maintain their own copies of NASA datasets. However, migration to the commercial cloud also poses some significant challenges, including (1) managing costs under a “pay-as-you-go” model; (2) incompatibility with existing tools and data formats with object-based storage and network access; (3) vendor lock-in; (4) challenges with data access for workflows that mix on-premise and cloud computing; and (5) standardization for highly diverse data as is present in NASA’s data archive. I conclude with two examples of recent NASA activities showcasing capabilities enabled by the commercial cloud: An interactive analysis and development platform for analyzing airborne imaging spectroscopy data, and a new collection of tools and services for data discovery, analysis, publication, and data-driven storytelling (Visualization, Exploration, and Data Analysis, VEDA).

Alexey N Shiklomanov↗

Fast and Flexible Multivariate Time Series Subsequence Search

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which often contain several gigabytes of data. Surprisingly, research on MTS search is very limited. Most of the existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two algorithms to solve this problem (1) a List Based Search (LBS) algorithm which uses sorted lists for indexing, and (2) a R*-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences. Both algorithms guarantee that all matching patterns within the specified thresholds will be returned (no false dismissals). The very few false alarms can be removed by a post-processing step. Since our framework is also capable of Univariate Time-Series (UTS) subsequence search, we first demonstrate the efficiency of our algorithms on several UTS datasets previously used in the literature. We follow this up with experiments using two large MTS databases from the aviation domain, each containing several millions of observations. Both these tests show that our algorithms have very high prune rates (>99%) thus needing actual disk access for only less than 1% of the observations. To the best of our knowledge, MTS subsequence search has never been attempted on datasets of the size we have used in this paper.

Bhaduri, Kanishka↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

Experiments on Supervised Learning Algorithms for Text Categorization

Modern information society is facing the challenge of handling massive volume of online documents, news, intelligence reports, and so on. How to use the information accurately and in a timely manner becomes a major concern in many areas. While the general information may also include images and voice, we focus on the categorization of text data in this paper. We provide a brief overview of the information processing flow for text categorization, and discuss two supervised learning algorithms, viz., support vector machines (SVM) and partial least squares (PLS), which have been successfully applied in other domains, e.g., fault diagnosis [9]. While SVM has been well explored for binary classification and was reported as an efficient algorithm for text categorization, PLS has not yet been applied to text categorization. Our experiments are conducted on three data sets: Reuter's- 21578 dataset about corporate mergers and data acquisitions (ACQ), WebKB and the 20-Newsgroups. Results show that the performance of PLS is comparable to SVM in text categorization. A major drawback of SVM for multi-class categorization is that it requires a voting scheme based on the results of pair-wise classification. PLS does not have this drawback and could be a better candidate for multi-class text categorization.

Namburu, Setu Madhavi↗

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS↗

Bundle Data Approach at GES DISC Targeting Natural Hazards

Severe natural phenomena such as hurricane, volcano, blizzard, flood and drought have the potential to cause immeasurable property damages, great socioeconomic impact, and tragic loss of human life. From searching to assessing the Big, i.e., massive and heterogeneous scientific data (particularly, satellite and model products) in order to investigate those natural hazards, it has, however, become a daunting task for Earth scientists and applications researchers, especially during recent decades. The NASA Goddard Earth Sciences Data and Information Service Center (GES DISC) has served Big Earth science data, and the pertinent valuable information and services to the aforementioned users of diverse communities for years. In order to help and guide our users to online readily (i.e., with a minimum effort) acquire their requested data from our enormous resource at GES DISC for studying their targeted hazard event, we have thus initiated a Bundle Data approach in 2014, first targeting the hurricane event topic. We have recently worked on new topics such as volcano and blizzard. The bundle data of a specific hazard event is basically a sophisticated integrated data package consisting of a series of proper datasets containing a group of relevant (knowledge--based) data variables readily accessible to users via a system-prearranged table linking those data variables to the proper datasets (URLs). This online approach has been developed by utilizing a few existing data services such as Mirador as search engine; Giovanni for visualization; and OPeNDAP for data access, etc. The online Data Cookbook site at GES DISC is the current host for the bundle data. We are now also planning on developing an Automated Virtual Collection Framework that shall eventually accommodate the bundle data, as well as further improve our management in Big Data.

GES DISC↗

Global Archaeal Diversity Revealed Through Massive Data Integration: Uncovering Just Tip of Iceberg

The domain of Archaea has gathered significant interest for its ecological and biotechnological potential and its role in helping us to understand the evolutionary history of Eukaryotes. In comparison to the bacterial domain, the number of adequately described members in Archaea is relatively low, with less than 1000 species described. It is not clear whether this is solely due to the cultivation difficulty of its members or, indeed, the domain is characterized by evolutionary constraints that keep the number of species relatively low. Based on molecular evidence that bypasses the difficulties of formal cultivation and characterization, several novel clades have been proposed, enabling insights into their metabolism and physiology. Given the extent of global sampling and sequencing efforts, it is now possible and meaningful to question the magnitude of global archaeal diversity based on molecular evidence. To do so, we extracted all sequences classified as Archaea from 500 thousand amplicon samples available in public repositories. After processing through our highly conservative pipeline, we named this comprehensive resource the ‘Global Archaea Diversity’ (GAD), which encompassed nearly 3 million molecular species clusters at 97% similarity, and organized it into over 500 thousand genera and nearly 100 thousand families. Saline environments have contributed the most to the novel taxa of this previously unseen diversity. The majority of those 16S rRNA gene sequence fragments were verified by matches in metagenomic datasets from IMG/M. These findings reveal a vast and previously overlooked diversity within the Archaea, offering insights into their ecological roles and evolutionary importance while establishing a foundation for the future study and characterization of this intriguing domain of life.

59 BASIC BIOLOGICAL SCIENCES↗

Investigation of low-energy particle remnants in high-energy collisions at the LHC with a skipper-CCD detector

We deployed the Mobile Skipper Testing Apparatus ∼33 m away from the Compact Muon Solenoid collision point, the first skipper-CCD detector probing low-energy particles produced in high-energy collisions at the Large Hadron Collider. In this work, we search for beam-related events using data collected in 2024 during beam-on and beam-off periods. The dataset corresponds to integrated luminosities of 113.3 fb −1 and 1.54 nb −1 for the proton-proton and Pb-Pb collision periods, respectively. We report observed event rates in a model-independent framework across two ionization regions: ≤ 20⁢𝑒 − and > 20⁢𝑒 − . For the low-energy region, we perform a likelihood analysis to test the null hypothesis of no beam-correlated signal. We found no significant correlation during proton-proton and Pb-Pb collisions. For the high-energy region, we present the energy spectra for both collision periods and compare event rates for images with and without luminosity. We observe a slight increase in the event rate following the Pb-Pb collisions, coinciding with a rise in the single-electron rate, which will be investigated in future work. Using the low-energy proton-proton results, we place 95% confidence level constraints on the mass-millicharge parameter space of millicharged particles. Overall, the results in this work demonstrate the viability of skipper-CCD technology to explore new physics at high-energy colliders and motivate future searches with more massive detectors.

Cervantes-Vergara, Brenda A. [Fermi National Accel↗