Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data

The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Advanced Computing, Data Science, and Artificial Intelligence Research Opportunities for Energy-Focused Transportation Science

The Energy Efficient Mobility Systems (EEMS) technology landscape is complex and rapidly evolving, which provides both tremendous opportunities and formidable challenges. Significant alterations to the mobility landscape are underway due to the advent of vehicle and infrastructure connectivity, autonomous driving, and rapid passenger- and freight-vehicle electrification. Advanced computing will play an increasingly important role in enabling the EEMS program to understand and identify the most important levers to improve the energy productivity of future integrated mobility systems. It is also driving new approaches to mobility and the research to unlock an affordable, efficient, safe, and accessible transportation future. Driving much of this change is the collection, analysis, and strategic use of massive amounts of diverse, complex data from infrastructure and vehicles with on-board sensors and data storage and transmission capabilities. Diverse and representative data are key to implementing approaches to maximize mobility energy productivity. While high-fidelity modeling of integrated transportation networks has strengthened our understanding of dynamic movement and behavior patterns, existing tools must be expanded beyond their current focus. This work necessitates data infrastructure investments (e.g., secure-streaming data platforms driven by ubiquitous sensors and video analytics) as well as investments in critical capabilities for large-scale automated analysis and organization using modern machine learning, statistics, and artificial intelligence. Other chief needs include agile, large-scale storage that can be quickly searched and queried for relevant data to support validation and model development, data-sharing agreements, and formatting standards for key data types. The future of public transit must be explored in greater detail, research must inform design, and opportunities must be identified for improving the mobility productivity of public transit in both urban and rural America.

33 ADVANCED PROPULSION SYSTEMS↗

PAS: Privacy Algorithms in Systems

Today we face an explosion of data generation, ranging from health monitoring to national security infrastructure systems. More and more systems are connected to the Internet that collects data at regular time intervals. These systems share data and use machine learning methods for intelligent decisions, which resulted in numerous real-world applications (e.g., autonomous vehicles, recommendation systems, and heart-rate monitoring) that have benefited from it. However, these approaches are prone to identity thief and other privacy related cyber-security attacks. So, how can data privacy be protected efficiently in these scenarios? More dedicated efforts are needed to propose the integration of privacy techniques into existing systems and develop more advanced privacy techniques to address the complex challenges of multi-system connectivity and data fusion. Therefore, we have introduced Privacy Algorithms in Systems (PAS) at CIKM which provides a venue to gather academic researchers and industry researchers/practitioners to present their research in an effort to advance the frontier of this critical direction of privacy algorithms in systems.

Kotevska, Olivera↗

Metadata Standards for the NSE: Extended Field Standards

This standard presents a set of optional metadata fields for managed digital objects within the Nuclear Security Enterprise (NSE) and provides a deeper look at data representation in metadata by looking at the representation of 1) Records Management required metadata, and 2) common representations of technical/scientific data. Metadata standardization is a critical enabler for effectively sharing data, documents, and other digital objects between NSE sites, and for tracing the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a complementary field standard, recommending an optional set of fields that should be uniformly built for all managed digital objects within the NSE. This document specifically focuses on extending the shared discovery layer defined in the first white paper by introducing additional descriptive and data representation fields that improve cross-site search and interpretation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Source Term Analysis of Xenon (STAX): An effort focused on differentiating man-made isotope production from nuclear explosions via stack monitoring

An overview of the hardware and software developed for the Source Term Analysis of Xenon (STAX) project is presented which includes the data collection from two stack monitoring systems installed at medical isotope production facilities, infrastructure to transfer data to a central repository, and methods for sharing data from the repository with users. STAX is an experiment to collect radioxenon emission data from industrial nuclear facilities with the goal of developing a better understanding of the global radioxenon background and the effect industrial radioxenon releases have on nuclear explosion monitoring. The final goal of this work is to utilize collected data along with atmospheric transport modeling to calculate the contribution of a peak or set of peaks detected by the International Monitoring System (IMS) to provide desired discriminating information to the International Data Centre (IDC) and National Data Centers (NDCs). Types of data received from the STAX equipment are shown and collected data was used for a case study to predict radioxenon concentrations at two IMS stations closest to the Institute for RadioElements (IRE) in Belgium. The initial evaluation of results indicate that the data is very valuable to the nuclear explosion monitoring community.

07 ISOTOPE AND RADIATION SOURCES↗

Leveraging History to Predict Infrequent Abnormal Transfers in Distributed Workflows

Scientific computing heavily relies on data shared by the community, especially in distributed data-intensive applications. This research focuses on predicting slow connections that create bottlenecks in distributed workflows. In this study, we analyze network traffic logs collected between January 2021 and August 2022 at the National Energy Research Scientific Computing Center (NERSC). Based on the observed patterns, we define a set of features primarily based on history for identifying low-performing data transfers. Typically, there are far fewer slow connections on well-maintained networks, which creates difficulty in learning to identify these abnormally slow connections from the normal ones. We devise several stratified sampling techniques to address the class-imbalance challenge and study how they affect the machine learning approaches. Our tests show that a relatively simple technique that undersamples the normal cases to balance the number of samples in two classes (normal and slow) is very effective for model training. This model predicts slow connections with an F1 score of 0.926.

97 MATHEMATICS AND COMPUTING↗

Integration of evidence across human and model organism studies: A meeting report

The National Institute on Drug Abuse and Joint Institute for Biological Sciences at the Oak Ridge National Laboratory hosted a meeting attended by a diverse group of scientists with expertise in substance use disorders (SUDs), computational biology, and FAIR (Findability, Accessibility, Interoperability, and Reusability) data sharing. The meeting's objective was to discuss and evaluate better strategies to integrate genetic, epigenetic, and 'omics data across human and model organisms to achieve deeper mechanistic insight into SUDs. Specific topics were to (a) evaluate the current state of substance use genetics and genomics research and fundamental gaps, (b) identify opportunities and challenges of integration and sharing across species and data types, (c) identify current tools and resources for integration of genetic, epigenetic, and phenotypic data, (d) discuss steps and impediment related to data integration, and (e) outline future steps to support more effective collaboration—particularly between animal model research communities and human genetics and clinical research teams. This review summarizes key facets of this catalytic discussion with a focus on new opportunities and gaps in resources and knowledge on SUDs.

59 BASIC BIOLOGICAL SCIENCES↗

Positron emission tomography harmonization in the Alzheimer's Disease Neuroimaging Initiative: A scalable and rigorous approach to multisite amyloid and tau quantification

Abstract INTRODUCTION A key goal of the Alzheimer's Disease NeuroImaging Initiative (ADNI) positron emission tomography (PET) Core is to harmonize quantification of β‐amyloid (Aβ) and tau PET image data across multiple scanners and tracers. METHODS We developed an analysis pipeline (Berkeley PET Imaging Pipeline, B‐PIP) for ADNI Aβ and tau PET images and applied it to PET data from other multisite studies. Steps include image pre‐processing, refacing, magnetic resonance imaging (MRI)/PET co‐registration, visual quality control (QC), quantification of tracer uptake, and standardization of Aβ and tau standardized uptake value ratios (SUVrs) across tracers. RESULTS Measurements from 10,105 cross‐sectional and longitudinal Aβ and tau PET scans acquired in several studies between 2010 and 2024 can be processed, harmonized, and directly merged across tracers and cohorts. DISCUSSION The B‐PIP developed in ADNI is a scalable image harmonization approach used in several observational studies and clinical trials that facilitates rigorous Aβ and tau PET quantification and data sharing. Highlights Quantitative results from ADNI Aβ and tau PET data are generated using a rigorous, scalable image processing pipeline This pipeline has been applied to PET data from several other large, multisite studies and trials Quantitative outcomes are harmonizable across studies and are shared with the scientific community

Neurosciences & Neurology↗

Comparison of Deterministic and Statistical Models for Water Quality Compliance Forecasting in the San Joaquin River Basin, California

Model selection for water quality forecasting depends on many factors including analyst expertise and cost, stakeholder involvement and expected performance. Water quality forecasting in arid river basins is especially challenging given the importance of protecting beneficial uses in these environments and the livelihood of agricultural communities. In the agriculture-dominated San Joaquin River Basin of California, real-time salinity management (RTSM) is a state-sanctioned program that helps to maximize allowable salt export while protecting existing basin beneficial uses of water supply. The RTSM strategy supplants the federal total maximum daily load (TMDL) approach that could impose fines associated with exceedances of monthly and annual salt load allocations of up to $1 million per year based on average year hydrology and salt load export limits. The essential components of the current program include the establishment of telemetered sensor networks, a web-based information system for sharing data, a basin-scale salt load assimilative capacity forecasting model and institutional entities tasked with performing weekly forecasts of river salt assimilative capacity and scheduling west-side drainage export of salt loads. Web-based information portals have been developed to share model input data and salt assimilative capacity forecasts together with increasing stakeholder awareness and involvement in water quality resource management activities in the river basin. Two modeling approaches have been developed simultaneously. The first relies on a statistical analysis of the relationship between flow and salt concentration at three compliance monitoring sites and the use of these regression relationships for forecasting. The second salt load forecasting approach is a customized application of the Watershed Analysis Risk Management Framework (WARMF), a watershed water quality simulation model that has been configured to estimate daily river salt assimilative capacity and to provide decision support for real-time salinity management at the watershed level. Analysis of the results from both model-based forecasting approaches over a period of five years shows that the regression-based forecasting model, run daily Monday to Friday each week, provided marginally better performance. However, the regression-based forecasting model assumes the same general relationship between flow and salinity which breaks down during extreme weather events such as droughts when water allocation cutbacks among stakeholders are not evenly distributed across the basin. A recent test case shows the utility of both models in dealing with an exceedance event at one compliance monitoring site recently introduced in 2020.

54 ENVIRONMENTAL SCIENCES↗

Electronic Visualization Laboratory's 50th Anniversary Retrospective: Look to the Future, Build on the Past

September 2023 marks the 50th anniversary of the Electronic Visualization Laboratory (EVL) at University of Illinois Chicago (UIC). EVL's introduction of the CAVE Automatic Virtual Environment in 1992, the first widely replicated, projection-based, walk-in, virtual-reality (VR) system in the world, put EVL at the forefront of collaborative, immersive data exploration and analytics. However, the journey did not begin then. Since its founding in 1973, EVL has been developing tools and techniques for real-time, interactive visualizations—pillars of VR. But EVL's culture is also relevant to its successes, as it has always been an interdisciplinary lab that fosters teamwork, where each person's expertise contributes to the development of the necessary tools, hardware, system software, applications, and human interface models to solve problems. Over the years, as multidisciplinary collaborations evolved and advanced scientific instruments and data resources were distributed globally, the need to access and share data and visualizations while working with colleagues, local and remote, synchronous and asynchronous, also became important fields of study. This paper is a retrospective of EVL's past 50 years that surveys the many networked, immersive, collaborative visualization and VR systems and applications it developed and deployed, as well as lessons learned and future plans.

Johnson, Andrew E.↗

A Functional All-Hazard Approach to Critical Infrastructure Dependency Analysis

The critical infrastructures protection landscape is a vast and varied pattern of independent, but interconnected infrastructure systems that are essential to the function of our modern society. The U.S. policy on critical infrastructure protection has been continually evolving since the “President’s Commission on Critical Infrastructure Protection” was published in 1997. In response to these policies, federal, state, and local governments, along with research institutions, have invested a substantial amount of time and effort into identifying and analyzing critical infrastructure, their functions, and dependencies/interdependencies to better understand their vulnerabilities. To date, the ability to assess vulnerabilities, resiliency, and priorities for protecting interdependent critical infrastructure systems from an all-hazards perspective remains a difficult problem. In this paper we introduce the All-Hazards Analysis (AHA) methodology, which provides an integrated functional basis across infrastructure systems, through the implementation of a common language and a scalable level of decomposition to effectively evaluate the resilience of interconnected infrastructure systems. AHA models infrastructure systems as directed multidimensional graphs, which enable the evaluation of cross-sector interdependencies prior to, during, and after disruptive events. Finally, and by design, AHA enables the cross linking of data taxonomies to enable more effective data sharing, such as the National Critical Functions (NCF) and Infrastructure Data Taxonomy (IDT).

02 PETROLEUM↗

Livewire Data Platform: File Standards Version 1.0

This technical document is a user guide to help users of the Livewire Data Platform understand the standards and requirements for storing and sharing data on the Livewire Data Platform.

97 MATHEMATICS AND COMPUTING↗

Data Quality Assessment of Optiwatt Vehicle Telematics Data

In October 2024, the Idaho National Laboratory (INL) received data from Optiwatt (Compass Global, Inc.) describing the driving and charging behavior of electric vehicle (EV) drivers. The data shared had been collected from approximately 10,000 vehicles and included vehicle specifications, driving information like odometer readings at the beginning and end of origin-destination pairs (i.e., trips with identification of home for trip start and end for Tesla vehicles), and charging information such as charging energy consumed per charge session and if the charge occurred at home. The vehicle data were provided from 9 EV makes and 18 EV models, with production years ranging from 2012–2024, but more than 9,500 of the vehicles were Tesla EVs. The data includes more than six million trips and more than three million charging events that occurred between June 2023 to Aug 2024 and collected from California and the Eastern United States. The purpose of this report is to review the quality of the data received from Optiwatt and the feedback INL received from Optiwatt after data concerns were shared with them.

33 - ADVANCED PROPULSION SYSTEMS↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

WA-Omic_LA.1.0 - Quantitative Lipidomics, Metabolomics, and Sequencing (16S/ITS) Publication Data DOI Package

Corresponding Data Publication: "Rapid remodeling of the soil lipidome in response to a drying-rewetting event." This study reveals specific changes in lipids and metabolites that are indicative of stress adaptation, substrate use, and cellular recovery during soil drying and subsequent rewetting. Drought induced nutrient limitation was reflected in the lipidome and polar metabalome, both of which rapidly shifted (within hours) upon rewet. Reduced nutrient access in dry soil caused the replacement of glycerophospholipids with phosphorus-free lipids and impeded resource-expensive osmolyte accumulation. Elevated levels of ceramides and lipids with long chain polyunsaturated fatty acids, in dry soil suggests that lipids play an important role in fungal drought tolerance. Increasing abundance of bacterial glycerophospholipids and triacylglycerols with fatty acids typical of bacteria and polar metabolites suggest metabolic recovery in representative bacteria once the environmental conditions are conducive for growth. These results underscore the importance of the soil lipidome as a robust indicator of microbial community responses, especially at the short time scales of cell-environment reactions. Data package contents reported here are the first version and contain pre- and post-processed data acquisition and subsequent downstream analysis files using various data source instrument method techniques and Mass Spectroscopy (MS) EMSL capabilities. This publication data package DOI is a comprehensive high-throughput multi-omics data lifecycle collection containing processed data method metadata. Support files include additional data download “Read Me” file containing data descriptor information and data source application ontologies (see data dictionary). Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. For increased data availability and interoperability, GC-MS/LC-MS mass spectrometry datasets (Thermo .raw ) were deposited at the MassIVE database repository under the related data accession MSV000086931 and can be accessed by using the API. Statistical data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location.

Amplicon sequencing 16S ITS LC-MS/MS lipidomics mu↗

WA-Omic_LA.1.0 - Quantitative Lipidomics, Metabolomics, and Sequencing (16S/ITS) Publication Data DOI Package

Corresponding Data Publication: "Rapid remodeling of the soil lipidome in response to a drying-rewetting event." This study reveals specific changes in lipids and metabolites that are indicative of stress adaptation, substrate use, and cellular recovery during soil drying and subsequent rewetting. Drought induced nutrient limitation was reflected in the lipidome and polar metabalome, both of which rapidly shifted (within hours) upon rewet. Reduced nutrient access in dry soil caused the replacement of glycerophospholipids with phosphorus-free lipids and impeded resource-expensive osmolyte accumulation. Elevated levels of ceramides and lipids with long chain polyunsaturated fatty acids, in dry soil suggests that lipids play an important role in fungal drought tolerance. Increasing abundance of bacterial glycerophospholipids and triacylglycerols with fatty acids typical of bacteria and polar metabolites suggest metabolic recovery in representative bacteria once the environmental conditions are conducive for growth. These results underscore the importance of the soil lipidome as a robust indicator of microbial community responses, especially at the short time scales of cell-environment reactions. Data package contents reported here are the first version and contain pre- and post-processed data acquisition and subsequent downstream analysis files using various data source instrument method techniques and Mass Spectroscopy (MS) EMSL capabilities. This publication data package DOI is a comprehensive high-throughput multi-omics data lifecycle collection containing processed data method metadata. Support files include additional data download “Read Me” file containing data descriptor information and data source application ontologies (see data dictionary). Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. For increased data availability and interoperability, GC-MS/LC-MS mass spectrometry datasets (Thermo .raw ) were deposited at the MassIVE database repository under the related data accession MSV000086931 and can be accessed by using the API. Statistical data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location.

Amplicon sequencing 16S ITS LC-MS/MS lipidomics mu↗

WA-Omic_LA.1.0 - Quantitative Lipidomics, Metabolomics, and Sequencing (16S/ITS) Publication Data DOI Package

Corresponding Data Publication: "Rapid remodeling of the soil lipidome in response to a drying-rewetting event." This study reveals specific changes in lipids and metabolites that are indicative of stress adaptation, substrate use, and cellular recovery during soil drying and subsequent rewetting. Drought induced nutrient limitation was reflected in the lipidome and polar metabalome, both of which rapidly shifted (within hours) upon rewet. Reduced nutrient access in dry soil caused the replacement of glycerophospholipids with phosphorus-free lipids and impeded resource-expensive osmolyte accumulation. Elevated levels of ceramides and lipids with long chain polyunsaturated fatty acids, in dry soil suggests that lipids play an important role in fungal drought tolerance. Increasing abundance of bacterial glycerophospholipids and triacylglycerols with fatty acids typical of bacteria and polar metabolites suggest metabolic recovery in representative bacteria once the environmental conditions are conducive for growth. These results underscore the importance of the soil lipidome as a robust indicator of microbial community responses, especially at the short time scales of cell-environment reactions. Data package contents reported here are the first version and contain pre- and post-processed data acquisition and subsequent downstream analysis files using various data source instrument method techniques and Mass Spectroscopy (MS) EMSL capabilities. This publication data package DOI is a comprehensive high-throughput multi-omics data lifecycle collection containing processed data method metadata. Support files include additional data download “Read Me” file containing data descriptor information and data source application ontologies (see data dictionary). Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. For increased data availability and interoperability, GC-MS/LC-MS mass spectrometry datasets (Thermo .raw ) were deposited at the MassIVE database repository under the related data accession MSV000086931 and can be accessed by using the API. Statistical data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location.

Amplicon sequencing 16S ITS LC-MS/MS lipidomics mu↗

dCache project status and update

The dCache project delivers an open-source, massively scalable, distributed storage system deployed internationally to satisfy today’s scientists’ ever-demanding storage requirements. Its multifaceted approach supports different use cases with the same storage, from high throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and longterm data persistence on tertiary storage. Even though dCache was initially developed for HEP experiments, today, it is used by various scientific communities, including astrophysics, biomed, and life science, each with their specific requirements. To match the needs of these new communities and keep up with the scaling demands of existing experiments, dCache is permanently evolving. With this contribution, we would like to highlight the recent developments in dCache regarding integration with CERN Tape Archive (CTA), advanced metadata handling, token-based authorization support, bulk API for QoS transitions, REST API to control interaction with the tape system, and future development directions.

Mkrtchyan, Tigran [DESY]↗