Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “software compatibility”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Accelerating GNNs on GPU Sparse Tensor Cores through N:M Sparsity-Oriented Graph Reordering

Recent advancements in GPU hardware support have introduced the capability to leverage N:M sparse patterns for substantial performance gains. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to such sparse patterns. In this paper, we propose a novel graph reordering algorithm, the first of its kind, to reshape irregular graph data into the N:M structured sparse pattern at the tile level, allowing linear-algebra-based graph operations in GNNs to benefit from the N:M sparse hardware. The optimization is lossless, maintaining the accuracy of GNN. It can remove 98-100\% violations of the N:M sparse patterns at the vector level, and increase the proportion of conforming graphs in SuiteSparse collection from 5-9\% to 88.7-93.5\%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (2.3X -- 7.5X on average) and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average).

artificial intelligence, graph neural networks↗

Agilent AgileBioFoundry CRADA (Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOF-MS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Therefore, our accomplishment in this CRADA improved the efficiency and accuracy of strain testing by developing and implementing fast analytical methods, robust processing tools, and software for predicting the best methods for UHPLC analysis.

97 MATHEMATICS AND COMPUTING↗

Development and implementation of high-throughput proteomic and metabolomics assays by using advanced chromatographic and mass spectrometric systems (CRADA Final Report)

The mission of this CRADA with Agilent was to couple powerful MS platforms (QQQ, IM-QTOFMS) with Agilent’s novel Ultra-High-Performance Liquid Chromatography (UHPLC) fast metabolomic workflows and perform ABF Machine Learning (ML) to generated datasets. Agilent transferred UHPLC methods to PNNL and LBNL and methods were implemented and demonstrated in both labs, achieving total acquisition times of < 10 min. Metabolites analyzed using Agilent’s shared methods included metabolites from central carbon metabolism, common across hosts, and metabolites unique to engineered strains. Standards were acquired in an UHPLC-Drift Tube Ion Mobility Mass Spectrometer (DTIMS) system for the first time within the context of ABF and methods were optimized based on Agilent’s protocols. Samples from ABF hosts Pseudomonas putida, Aspergillus pseudoterreus, Aspergillus niger and Rhodosporidium toruloides were analyzed using the UHPLC-DTIMS platform for a total of 276 runs. A data analysis workflow compatible with the Experimental Data Depot (EDD) and completely shareable was developed for the acquired UHPLC-DTIMS data. Samples were analyzed using a Data Independent Acquisition Approach (DIA), which for most of the standards provided more transitions therefore increasing detection confidence. Using the data acquired by PNNL, LBNL, and Agilent’s specifications from previous ML projects, SNL applied an ensemble ML strategy to pick the best performing model for automated LC-method selection. Finally, with the contribution of the participant labs and Agilent, SNL developed an Automated Method Selection (AMS) software tool to predict the best liquid chromatography method for analysis of any new molecules of interest. Samples with novel pathways and new metabolite targets of interest are generated at a high pace in the ABF. Overall, the project advanced rapid metabolomics by combining liquid chromatography, ion mobility spectrometry, and data-independent mass spectrometry with machine learning. This multidimensional approach uses retention time, collision cross-section, precursor mass, and fragment-ion information to distinguish chemically similar metabolites that can be difficult to resolve using conventional liquid- or gas-chromatography methods. The resulting workflow also provided automated metabolite-identification error estimates, addressing a recognized need for statistical confidence measures in metabolomics.

Petzold, Christopher [Lawrence Berkeley National L↗

Integration of RNTuple in ATLAS Athena

After using ROOT’s TTree I/O subsystem for over two decades and storing more than an exabyte of compressed High Energy Physics (HEP) data, advances in technology have motivated a complete redesign, RNTuple, which breaks backward-compatibility to take better advantage of these storage options. The RNTuple I/O subsystem has been designed to address performance bottlenecks and other shortcomings of TTree. Specifically, RNTuple comes with an updated, more compact binary data format that can be stored both in ROOT files and natively in object stores. It is designed for modern storage hardware (e.g. high-throughput low-latency NVMe SSDs), and provides robust and easy to use interfaces. The binary format of RNTuple is scheduled to become production grade in 2024, and recently has become mature enough to start exploring the integration into software used by HEP experiments. In this contribution, we discuss the developments to support the features as required by the ATLAS analysis Event Data Model (EDM) in RNTuple, which will enable its integration into the Athena software framework. With these developments in place, we evaluate the performance of the current most recent versions of RNTuple-based ATLAS data sets and compare this to that of TTree.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

An automated platform for in situ serial crystallography at room temperature

Direct observation of functional motions in protein structures is highly desirable for understanding how these nanomachineries of life operate at the molecular level. Because cryogenic temperatures are non-physiological and may prohibit or even alter protein structural dynamics, it is necessary to develop robust X-ray diffraction methods that enable routine data collection at room temperature. We recently reported a crystal-on-crystal device to facilitate in situ diffraction of protein crystals at room temperature devoid of any sample manipulation. Here an automated serial crystallography platform based on this crystal-on-crystal technology is presented. A hardware and software prototype has been implemented, and protocols have been established that allow users to image, recognize and rank hundreds to thousands of protein crystals grown on a chip in optical scanning mode prior to serial introduction of these crystals to an X-ray beam in a programmable and high-throughput manner. This platform has been tested extensively using fragile protein crystals. We demonstrate that with affordable sample consumption, this in situ serial crystallography technology could give rise to room-temperature protein structures of higher resolution and superior map quality for those protein crystals that encounter difficulties during freezing. This serial data collection platform is compatible with both monochromatic oscillation and Laue methods for X-ray diffraction and presents a widely applicable approach for static and dynamic crystallographic studies at room temperature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Redis-Based Streaming Architecture for Accelerator Beam Instrumentation DAQ Systems

The Fermilab Acceleraor Division, Beam Instrumentation Department, is always adopting modern and current software methodologies for complex DAQ architectures. This paper highlights the Redis Adapter (RA) as the key software component enabling high performance, modular communication between digitizers and distributed control systems by leveraging Redis and containerization. The RA provides a unified, efficient interface between Redis based data streams and consumer systems. In the legacy architecture, digitized data flowed through the custom, UDP based Distributed Data Communication Protocol in the middle layer. In the current system, DDCP remains the ingestion path, while the RA serves as the decoupling layer. The proposed system replaces old VME digitizers with a SOM-based digitizer that communicates with Redis using the RA. The RA acts as both a performance-critical bridge and a protocol-agnostic adapter, ensuring compatibility with legacy control frameworks while enabling future scalability and modularity. This restructuring of the middle layer also helps the system achieve high throughput, reduce latency, and simplify the data path. Finally, we will demonstrate how RA is utilized in our two core products to deliver both legacy compatibility and future flexibility.

Joshi, S. [Fermilab]↗

CPU and memory efficient coherent mode decomposition for the partially coherent x-ray simulations

In this work, the method of the Coherent Mode Decomposition (CMD) is applied to numerical wave propagation calculations for partially-coherent X-rays, using the Fourier optics and compatible methods. Its CPU and memory efficiency is discussed in various cases of the wavefront at the source and the beam waist. With the absence of the quadratic phase terms, the required sampling density of the electric fields is effectively reduced. The problem size is thus moderate and the method is feasible to be implemented on a single-node CPU server. In other cases, the same argument holds with proper treatments of the quadratic phase terms. Tests on CMD and the modes propagation are done for the case of the Coherent Hard X-ray beamline of the National Synchrotron Light Source II, using the Synchrotron Radiation Workshop software. We observe a few hundred or less dominant decomposed modes that resemble the electric fields converge to the wavefront intensity at a high accuracy of over 99%.

36 MATERIALS SCIENCE↗

Geant4 based positron beam source (GPos) v1.0

GPos is a software that was created to determine the properties of positron beams resulting from the interaction of the LBNL BELLA center PetaWatt laser-driven plasma-capillary accelerated electron beam and the atoms of a thin solid target. GPos is written in C++, easily compiled with cmake and the spack package manager, which allows for multi-thread and MPI parallel computing. Its functions expand on the Geant4 toolkit library and allow for propagation of the modelled particles through vacuum drift distances with a focusing element (thin lens approximation). Users can change beam-foil-drift-lens parameters - to adapt GPos to other particle sources and infrastructures - in a simple input file. The code particle data output format, openPMD, which is compatible, for example, with the input of the ECP WarpX project code used to explore the physics of particle acceleration in plasmas. Using GPos in conjunction to WarpX allowed us to test various configurations for designing a high-quality and high-energy positron source at BELLA -required for us to address positron acceleration challenges in the development of future linear colliders. GPos can also be advantageous when tackling the physics of muon sources for future muon colliders as well as for the investigation of positron sources in lower energy regimes for applications like annihilation spectroscopy and astrophysical gamma-ray-bursts.

Pinto de Almeida Amorim, Ligia↗

Validation of standardized data formats and tools for ground-level particle-based gamma-ray observatories

Context. Ground-based γ-ray astronomy is still a rather young field of research, with strong historical connections to particle physics. This is why most observations are conducted by experiments with proprietary data and analysis software, as is usual in the particle physics field. However, in recent years, this paradigm has been slowly shifting toward the development and use of open-source data formats and tools, driven by upcoming observatories such as the Cherenkov Telescope Array (CTA). In this context, a community-driven, shared data format (the gamma-astro-data-format, or GADF) and analysis tools such as Gammapy and ctools have been developed. So far, these efforts have been led by the Imaging Atmospheric Cherenkov Telescope community, leaving out other types of ground-based γ-ray instruments. Aims. We aim to show that the data from ground particle arrays, such as the High-Altitude Water Cherenkov (HAWC) observatory, are also compatible with the GADF and can thus be fully analyzed using the related tools, in this case, Gammapy. Methods. We reproduced several published HAWC results using Gammapy and data products compliant with GADF standard. We also illustrate the capabilities of the shared format and tools by producing a joint fit of the Crab spectrum including data from six different γ-ray experiments. Results. We find excellent agreement with the reference results, a powerful confirmation of both the published results and the tools involved. Conclusions. The data from particle detector arrays such as the HAWC observatory can be adapted to the GADF and thus analyzed with Gammapy. A common data format and shared analysis tools allow multi-instrument joint analysis and effective data sharing. To emphasize this, a sample of Crab nebula event lists is made public with this paper. Because of the complementary nature of pointing and wide-field instruments, this synergy will be distinctly beneficial for the joint scientific exploitation of future observatories such as the Southern Wide-field Gamma-ray Observatory and CTA.

79 ASTRONOMY AND ASTROPHYSICS↗

The Importance of Freeze/Thaw Cycles on Lateral Transport in Ice-Wedge Polygons: Modeling Archive

This Modeling Archive is in support of an NGEE Arctic publication "The importance of freeze/thaw cycles on lateral transport in ice-wedge polygons". The dataset includes xml input/configuration files. These files are compatible with the ATS version 1.0 and higher. The mesh folder contains mesh files used for high-centered polygon (hcp) and low-centered polygon (lcp). The mesh files represent the transect of the polygonal tundra (Fig 1). Here we used two types of mesh with impermeable layer and without. The mesh with an impermeable layer corresponds to the synthetic permafrost (no freezeup case). The freeze up case uses mesh without impermeable layer to simulated frozen ground start at the same depth where the impermeable layer is for the no freezeup case. The meteorological data used drive the model saved in the "inputs" folder. The processed tracer flow rates are saved in the "tracer-flow-ratesCfolder. To plot figures 1 and 2, we used VisIt software. To plot all the flow rates, we used ipython notebook script. All required inputs are saved in "freezeup" and "no freezeup" folders. Each folder includes the corresponding "lcp" and "hcp" folders. The "freezeup" folder has also flat-centered polygon (fcp) case, low porosity "lpor", and low permeability "lper" cases. Included are *.xml, *.exo, *.h5, *pdf, *.ipynb, *.sh, *.py, and *.out files. NGEE Arctic Project Summary: The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

The Silver Lining

Clouds are shareable scientific instruments that create the potential for reproducibility by ensuring that all investigators have access to a common execution platform on which computational experiments can be repeated and compared. By virtue of the interface they present, they also lead to the creation of digital artifacts compatible with the cloud, such as images or orchestration templates, that go a long way-and sometimes all the way-to representing an experiment in a digital, repeatable form. In this article, I describe how we developed these natural advantages of clouds in the Chameleon testbed and argue that we should leverage them to create a digital research marketplace that would make repeating experiments as natural and viable part of research as sharing ideas via reading papers is today.

Cloud computing↗

Reaction Mechanism Generator v3.0: Advances in Automatic Mechanism Generation

In chemical kinetics research, kinetic models containing hundreds of species and tens of thousands of elementary reactions are commonly used to understand and predict the behavior of reactive chemical systems. Reaction Mechanism Generator (RMG) is a software suite developed to automatically generate such models by incorporating and extrapolating from a database of known thermochemical and kinetic parameters. Here, we present the recent version 3 release of RMG and highlight improvements since the previously published description of RMG v1.0. Most notably, RMG can now generate heterogeneous catalysis models in addition to the previously available gas- and liquid-phase capabilities. For model analysis, new methods for local and global uncertainty analysis have been implemented to supplement first-order sensitivity analysis. The RMG database of thermochemical and kinetic parameters has been significantly expanded to cover more types of chemistry. The present release includes parallelization for faster model generation and a new molecule isomorphism approach to improve computational performance. RMG has also been updated to use Python 3, ensuring compatibility with the latest cheminformatics and machine learning packages. Overall, RMG v3.0 includes many changes which improve the accuracy of the generated chemical mechanisms and allow for exploration of a wider range of chemical systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TEMPI: An Interposed MPI Library with Canonical Representation of MPI Datatypes [Poster]

TEMPI provides a transparent non-contiguous data-handling layer compatible with various MPIs. MPI Datatypes are a powerful abstraction for allowing an MPI implementation to operate on non-contiguous data. CUDA-aware MPI implementations must also manage transfer of such data between the host system and GPU. The non-unique and recursive nature of MPI datatypes mean that providing fast GPU handling is a challenge. The same noncontiguous pattern may be described in a variety of ways, all of which should be treated equivalently by an implementation. This work introduces a novel technique to do this for strided datatypes. Methods for transferring non-contiguous data between the CPU and GPU depends on the properties of the data layout. This work shows that a simple performance model can accurately select the fastest method. Unfortunately, the combination of MPI software and system hardware available may not provide sufficient performance. The contributions of this work are deployed on OLCF Summit through an interposer library which does not require privileged access to the system to use

97 MATHEMATICS AND COMPUTING↗

AmeriFlux FLUXNET-1F US-Cst Crossett Experimental Forest

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-Cst Crossett Experimental Forest. This is the FLUXNET version of the carbon flux data for the site US-Cst Crossett Experimental Forest produced by applying the standard ONEFlux (1F) software. Site Description - The study takes place in the Crossett Experimental Forest (CEF, 32°2′ N, 91°57′ W), which was established in 1934 with the objective of developing effective management protocols for loblolly and shortleaf pine forests in the Gulf Coastal Plain. Mean annual temperature and precipitation are 17.6 °C and 1,410 mm, respectively. Most of the soils on and near the CEF are silt loams (primarily Glossaquic Fragiudalfs). The tower is situated in the the old-growth pine management sector of the CEF, which emphasizes management for old-growth-like conditions, with the goal of emulating the structural and compositional attributes of historic forests from this region using a combination of timber harvesting, prescribed fire, and other tools. This includes periodic commercial logging of intermediate-sized loblolly pines to support further restoration treatments, such as prescribed fire, hardwood and exotic species competition control, and the increase of plant species more compatible with fire (e.g., shortleaf pine and understory grasses)

Novick, Kim↗

A metabolic modeling platform for the computation of microbial ecosystems in time and space (COMETS)

Genome-scale stoichiometric modeling of metabolism has become a standard systems biology tool for modeling cellular physiology and growth. Extensions of this approach are emerging as a valuable avenue for predicting, understanding and designing microbial communities. Computation of microbial ecosystems in time and space (COMETS) extends dynamic flux balance analysis to generate simulations of multiple microbial species in molecularly complex and spatially structured environments. Here we describe how to best use and apply the most recent version of COMETS, which incorporates a more accurate biophysical model of microbial biomass expansion upon growth, evolutionary dynamics and extracellular enzyme activity modules. In addition to a command-line option, COMETS includes user-friendly Python and MATLAB interfaces compatible with the well-established COBRA models and methods, as well as comprehensive documentation and tutorials. Overall, this protocol provides a detailed guideline for installing, testing and applying COMETS to different scenarios, generating simulations that take from a few minutes to several days to run, with broad applicability to microbial communities across biomes and scales.

59 BASIC BIOLOGICAL SCIENCES↗

Behind the Meter Storage for Electric Vehicle Charging, Electrochemical and Thermal Energy Storage, and Solar Photovoltaic

In response to the potentially large and irregular demand from EVs, along with changing load profiles from buildings with on-site generation, utilities are evaluating multiple options for managing dynamic loads, including time-of-use pricing, demand charges, battery storage, and curtailment of variable generation. Buildings, as well as commercial, public, and workplace EV charging operations, can use a combination of electrochemical battery storage and thermal energy storage coupled with on-site generation to manage energy costs as well as provide resiliency and reliability for EV charging and building energy loads. We are completing a behind the meter storage analysis that focuses on determining the optimal system designs and energy flows for thermal and electrochemical behind the meter storage with on-site solar photovoltaic (PV) generation enabling electric vehicle charging in various climates, building types, and utility rate structures. In completing this analysis, we have developed a tool that combines existing battery models via the System Advisor Model (SAM) and building modeling software via EnergyPlus into a single interface. This tool allows us to simulate a building with a detailed battery model to properly size the battery, thermal energy storage, and solar PV systems to maximize profit for the system owner. This also allows us to assess how the battery degrades under various supervisory control dispatch algorithms to control charging/discharging; we can also see how thermal energy storage is created and used to complement the battery to reduce thermal loads in the building. With this project, we can analyze new batteries that are designed specifically for energy storage, rather than designed to be extremely energy dense for electric vehicle applications, using battery lifetime models from other national labs and the existing SAM battery model, which has detailed lifetime and degradation parameters. We can also assess novel thermal storage technologies by integrating them into the whole building energy simulation program EnergyPlus. Because the model calls both SAM and EnergyPlus, required inputs need to be compatible for both models. These inputs include, on a high-level, the following: weather files, building and electric vehicle load profiles, electricity rate tariff information, and system cost information for the stationary battery, solar PV, and thermal storage system. The various buildings we are studying for this analysis are retail big-box grocery store, commercial office building, fleet vehicle depot and operations facility, multi-family residential, and electric vehicle charging station. For these different applications, the battery and thermal storage will be dispatched differently, and the various technologies are sized differently to optimize cost.

30 DIRECT ENERGY CONVERSION↗

From Modular ADMS to Plug-and-Play Ops: Distribution Grid Operations with Platform-Level Orchestration to Enable Ambitious App Hosting

The core function of the distribution grid is to provide electricity to consumers affordably, reliably, and securely. In pursuing these core objectives, distribution utilities are accountable to customers, regulators, and in some cases, shareholders. Other third parties such as aggregators and microgrids can also have a stake in the smooth operation of the grid. Each of these stakeholders has economic, business, and/or governance objectives that inform their expectations of the distribution grid. This multi-objective, multi-stakeholder environment creates tension that must be reconciled to successfully design and operate the distribution grid. Innovative companies are competing to bring high-tech solutions to electric utilities and their customers that address each of these objectives. Many developers of advanced distribution management systems (ADMS) and distributed energy resource management systems (DERMS) have adopted a modular architecture that allows grid operators to select functions and features according to their individual system needs. A modular platform also allows the solution provider to develop and integrate specific new product modules; however, the need to pursue multiple objectives with a fixed set of controllable devices makes integration expensive whether it is done at the product development stage or the deployment stage. This cost creates a significant barrier to adoption and can lengthen the product to market time of new solutions. To fundamentally address the complexity of system integration for distribution grid operations, the U.S. Department of Energy Office of Electricity has funded the GridAPPS-D project at PNNL, which streamlines integration by contributing to standards development, defining system architecture, applying advanced mathematics, and developing open-source software to demonstrate the concept of an open data-integration platform for distribution operations. The open data-integration platform concept enables system operators and solution providers to deploy ambitious, best-of-breed applications (or apps) without continually reengineering for integration. Ambitious apps developed by different solution providers will inevitably attempt to achieve different control objectives with the same set of controllable devices. If the open platform itself can resolve these conflicts in a way that achieves the best available outcomes for all apps, doesn’t restrict the ambitious design of apps, and ensures safe and secure operations, apps will be able to plug-and-play with the platform at the same time as other ambitious apps. In this paper, we describe a framework called App Deconfliction that empowers a platform to assign setpoints to controllable devices based on the values preferred by different apps (and even external stakeholder entities like customers or aggregators). The App Deconfliction framework is compatible with several methods for determining setpoint values. We present two methods based on game theory that provide a subtle built-in incentive structure for developers to adapt their apps to the fact that they will be operating in a moderated multi-app environment and to favor device setpoints that have the most effect on their objectives over those that have the least effect. Our simulation-based demonstrations have shown that game-theory-based deconfliction can lead to a 7% improvement in control space utilization compared to design-based methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES↗