Engineering PapersSearch

SEARCH · Engineering Papers

Results for “dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Astronaut Photography of the Earth: A Long-Term Dataset for Earth Systems Research, Applications, and Education

The NASA Earth observations dataset obtained by humans in orbit using handheld film and digital cameras is freely accessible to the global community through the online searchable database at https://eol.jsc.nasa.gov, and offers a useful compliment to traditional ground-commanded sensor data. The dataset includes imagery from the NASA Mercury (1961) through present-day International Space Station (ISS) programs, and currently totals over 2.6 million individual frames. Geographic coverage of the dataset includes land and oceans areas between approximately 52 degrees North and South latitudes, but is spatially and temporally discontinuous. The photographic dataset includes some significant impediments for immediate research, applied, and educational use: commercial RGB films and camera systems with overlapping bandpasses; use of different focal length lenses, unconstrained look angles, and variable spacecraft altitudes; and no native geolocation information. Such factors led to this dataset being underutilized by the community but recent advances in automated and semi-automated image geolocation, image feature classification, and web-based services are adding new value to the astronaut-acquired imagery. A coupled ground software and on-orbit hardware system for the ISS is in development for planned deployment in mid-2017; this system will capture camera pose information for each astronaut photograph to allow automated, full georegistration of the data. The ground system component of the system is currently in use to fully georeference imagery collected in response to International Disaster Charter activations, and the auto-registration procedures are being applied to the extensive historical database of imagery to add value for research and educational purposes. In parallel, machine learning techniques are being applied to automate feature identification and classification throughout the dataset, in order to build descriptive metadata that will improve search capabilities. It is expected that these value additions will increase interest and use of the dataset by the global community.

Stefanov, William L.

Aircraft Engine Run-to-Failure Dataset Under Real Flight Conditions for Prognostics and Diagnostics

A key enabler of intelligent maintenance systems is the ability to predict the remaining useful lifetime (RUL) of its components, i.e., prognostics. The development of data-driven prognostics models requires datasets with run-to-failure trajectories. However, large representative run-to-failure datasets are often unavailable in real applications because failures are rare in many safety-critical systems. To foster the development of prognostics methods, we develop a new realistic dataset of run-to-failure trajectories for a fleet of aircraft engines under real flight conditions. The dataset was generated with the Commercial Modular Aero-Propulsion System Simulation (CMAPSS) model developed at NASA. The damage propagation modelling used in this dataset builds on the modelling strategy from previous work and incorporates two new levels of fidelity. First, it considers real flight conditions as recorded on board of a commercial jet. Second, it extends the degradation modelling by relating the degradation process to its operation history. This dataset also provides the health respectively fault class. Therefore, besides its applicability to prognostics problems, the dataset can be used for fault diagnostics.

CMAPPS

Visual and Inertial Datasets for an eVTOL Aircraft Approach and Landing Scenario

A National Aeronautics and Space Administration (NASA) project developing computer vision algorithms for autonomous flight is producing real-world datasets with cameras mounted on aircraft. In related domains, such as autonomous driving, open datasets are key to innovation and advancement in computer vision and autonomous perception for future Advanced Air Mobility (AAM) operations. Few vision datasets, however, are publicly available in the aviation context. This paper introduces preliminary datasets containing several examples of approach and landing scenarios. The platform aircraft include a multirotor small unmanned aerial system (sUAS) and a crewed helicopter as surrogates for future electric vertical take-off and landing (eVTOL) aircraft. The dataset provides video imagery with associated inertial navigation system-global positioning system (INS-GPS) position and attitude estimates and other sensors. Surveyed locations of the visual features of the landing area are included. This dataset is the first to be released in an ongoing effort to collect and share large, diverse datasets relevant to autonomous aviation; community critique that can inform and improve future flight campaigns is welcome.

Nelson Brown

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING

Search for ultralight dark matter in the SuperMAG high-fidelity dataset

Ultralight dark matter, such as kinetically mixed dark-photon dark matter (DPDM) or axion-like-particle dark matter (axion DM), can source an oscillating magnetic-field signal at Earth’s surface. Previous work searched for this signal in a publicly available dataset of global magnetometer measurements maintained by the SuperMAG collaboration. This “low-fidelity” dataset reported measurements with a 1-min time resolution, allowing the search to set leading direct constraints on DPDM and axion DM with Compton frequencies f DM ≤ 1 / ( 1 min ) (corresponding to masses m DM ≤ 7 × 10 − 17 eV ). More recently, a dedicated experiment undertaken by the SNIPE Hunt collaboration has also searched for this same signal at higher frequencies f DM ≥ 0.5 Hz (or m DM ≥ 2 × 10 − 15 eV ). In this work, we search for this signal of ultralight DM in the SuperMAG “high-fidelity” dataset, which features a 1-sec time resolution, allowing us to probe the gap in parameter space between the low-fidelity dataset and the SNIPE Hunt experiment. The high-fidelity dataset exhibits lower geomagnetic noise than the low-fidelity dataset and features more data than the SNIPE Hunt experiment, making it a powerful probe of ultralight DM. Our search finds no robust DPDM or axion DM candidates. We set constraints on DPDM and axion DM parameter space for 10 − 3 Hz ≤ f DM ≤ 0.98 Hz (or 4 × 10 − 18 eV ≤ m DM ≤ 4 × 10 − 15 eV ). Our results are the leading direct constraints on both DPDM and axion DM in this mass range, and our DPDM constraint surpasses the leading astrophysical constraint in a narrow range around m A ′ ≈ 2 × 10 − 15 eV . Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Data Fusion for the Development of a Multimodal Freight Transload Facilities Dataset in the U.S.

To withstand the growing demand of commodity volume and its strain on the transportation infrastructure, it is necessary to identify the flow of commodities by route and mode. However, a national multimodal freight routing model does not exist for the U.S. The development of such model requires multiple building blocks, such as virtual representations of roadway, railway, and waterway networks, transload facilities (TFs), and access/egress links. Most of these blocks have a robust database in the U.S., except for the TFs. Here, this paper presents the fusion of dispersed and heterogeneous representations of multimodal TFs into a single, comprehensive, geospatial freight TF dataset. The TF dataset is derived from several sources, including the U.S. Army Corps of Engineers Master Docks Plus, the National Transportation Atlas Database, the Intermodal Association of North America, industry publications, and other public information. First, individual datasets were queried and reconciled. A geocoding/reverse geocoding process was applied to get the best street address and latitude/longitude location for each terminal. Then, duplicate terminals were identified by a fuzzy match algorithm based on terminal name and location, and removed. Validation was performed by visual inspection of random facilities. The main contributions of this work are: a publicly available version of the TF dataset, including facility location and multimodal transfer capability of 9,003 facilities, and an enterprise-version with the same facilities but including commodity handling capabilities. The main purpose of developing the TF dataset is to inform multimodal routing algorithms. The proposed TF dataset allows for credibly modeling the multimodal transfer of commodities within shipment routes.

Commodity Routing

U.S. Freight Transload Facilities Dataset

The U.S. Freight Transload Facilities Dataset provides location information (latitude, longitude, zip, city, county, state)for more than 9,000 facilities across 50 U.S. States where freight may be transferred between waterways, railways, and roadways. The dataset lists the known modes and available direction(s) for freight transfers at each facility as of 2024. The U.S. Freight Transload Facilities dataset was built by mining and fusing several public sources, such as the USACE Master Docks Plus, the USDOT National Transportation Atlas Database (NTAD), files from the Intermodal Association of North America (IANA), and the industry publication Bulk Transloader. The dataset constitutes a key piece of a multimodal freight transportation network and routing algorithm developed by USACE-ERDC. The dataset is shared as a .csv file. The dataset is published for research purposes and should not be considered exhaustive or authoritative.

Peterson, Steven [ORNL] (ORCID:0000000287672998)

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY

Heuristics for Relevancy Ranking of Earth Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management

Relevancy Ranking of Satellite Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management

Development of a Knowledge Graph for Dataset Discovery and Identification at a NASA Data Center

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) archives and distributes hundreds of Earth Science data collections to the public. These collections are used in research, resulting in the publication of thousands of scientific papers each year. As new users come to GES DISC for data, it is important for them to understand how prior research used the data. To help researchers, a knowledge graph (KG) was designed and implemented to connect publication citations with dataset metadata. The relationships created in the graph have the potential to allow the Web applications that utilize this information to directly connect the publication to the GES DISC datasets and services. These relationships are demonstrated using a web application prototype. In addition, the graph can also make connections between publications, datasets, and measurements based on the mentions of datasets and their attributes in the publications. To demonstrate this capability, a web application was created that takes the excerpt from the publication and returns a most likely dataset and measurement pairing, ranking the results based on how often these datasets and measurements were used in prior publications.

Nathaniel Crosby

Application of a Dataset-Publication Knowledge Graph for Improving Earth Science Data Search

Finding a dataset at a NASA data center that is the best fit for the researcher’s application presents a challenge, not only for a novice user but for an experienced one, due to the data complexity and a multitude of choices of the existing data. Users often search for the data based on the application they are interested in, their research domain, phenomena, research topic, etc. As existing dataset metadata may not cover these search terms, the user may not obtain the most relevant results for their purpose. This problem was addressed by leveraging the content of the titles and abstracts of the research papers that utilize NASA datasets. For this, features from the paper titles and abstracts were extracted, and then a knowledge graph (KG) was used to link these features to the datasets used in that paper. The search for the datasets was tested by querying this knowledge graph through various terms extracted from Earth Science ontologies such as Semantic Web for Earth and Environment Technology (SWEET), and it was shown that this KG search outperforms the existing search that exclusively queries the dataset metadata.

Kristina Stoyanova

Global Total Ozone Recovery Trends Attributed to Ozone-Depleting Substance (ODS) Changes Derived From Five Merged Ozone Datasets

We report on updated trends using different merged zonal mean total ozone datasets from satellite and ground-based observations for the period from 1979 to 2020. This work is an update of the trends reported in Weber et al. (2018) using the same datasets up to 2016. Merged datasets used in this study include NASA MOD v8.7 and NOAA Cohesive Data (COH) v8.6, both based on data from the series of Solar Backscatter Ultraviolet (SBUV), SBUV-2, and Ozone Mapping and Profiler Suite (OMPS) satellite instruments (1978–present), as well as the Global Ozone Monitoring Experiment (GOME)-type Total Ozone – Essential Climate Variable (GTO-ECV) and GOME-SCIAMACHY-GOME-2 (GSG) merged datasets (both 1995–present), mainly comprising satellite data from GOME, SCIAMACHY, OMI, GOME-2A, GOME-2B, and TROPOMI. The fifth dataset consists of the annual mean zonal mean data from ground-based measurements collected at the World Ozone and Ultraviolet Radiation Data Centre (WOUDC). Trends were determined by applying a multiple linear regression (MLR) to annual mean zonal mean data. The addition of 4 more years consolidated the fact that total ozone is indeed slowly recovering in both hemispheres as a result of phasing out ozone-depleting substances (ODSs) as mandated by the Montreal Protocol. The near-global (60° S–60° N) ODS-related ozone trend of the median of all datasets after 1995 was 0.4 ± 0.2 (2σ) %/decade, which is roughly a third of the decreasing rate of 1.5 ± 0.6 %/decade from 1978 until 1995. The ratio of decline and increase is nearly identical to that of the EESC (equivalent effective stratospheric chlorine or stratospheric halogen) change rates before and after 1995, confirming the success of the Montreal Protocol. The observed total ozone time series are also in very good agreement with the median of 17 chemistry climate models from CCMI-1 (Chemistry-Climate Model Initiative Phase 1) with current ODS and GHG (greenhouse gas) scenarios (REF-C2 scenario). The positive ODS-related trends in the Northern Hemisphere (NH) after 1995 are only obtained with a sufficient number of terms in the MLR accounting properly for dynamical ozone changes (Brewer–Dobson circulation, Arctic Oscillation (AO), and Antarctic Oscillation (AAO)). A standard MLR (limited to solar, Quasi-Biennial Oscillation (QBO), volcanic, and El Niño–Southern Oscillation (ENSO)) leads to zero trends, showing that the small positive ODS-related trends have been balanced by negative trend contributions from atmospheric dynamics, resulting in nearly constant total ozone levels since 2000.

Total Column Ozone Trends

Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) Datasets Released by NASA GES DISC and Their Applications for Air Quality

Nitrogen dioxide (NO2), a pervasive air pollutant, comes from vehicles, power plants, industrial emissions, and off-road sources such as construction or lawn and gardening equipment. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) curates many remote sensing datasets with NO2 retrievals, which have been utilized for air quality research and applications. The remotely-sensed datasets include those generated by the Ozone Monitoring Instrument (OMI) on the Aura satellite, the TROPOspheric Monitoring Instrument (TROPOMI) onboard the Copernicus Sentinel-5 Precursor (S5P), and the Ozone Mapping and Profiling Suite (OMPS) Nadir-Mapper (NM) instrument on the Suomi National Polar-orbiting Partnership (S- NPP). In collaboration with the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) project, the GES DISC recently released MINDS datasets. The NASA MEaSUREs MINDS project aims to develop long-term NO2 global data records by adapting a consistent retrieval algorithm to multiple instrument measurements. Long-term data records will be achieved by applying consistent retrieval approaches to multiple satellite instruments, including OMI (2004 - ); the Global Ozone Monitoring Experiment (GOME, 1995-2011) onboard the second European Remote Sensing satellite (ERS-2); the Scanning Imaging Spectrometer for Atmospheric Cartography (SCIAMACHY, 2002-2012) onboard the ENVIronmental SATellite (ENVISAT); GOME-2 on the Meteorological Operational satellites (MetOp-A and MetOp-B, 2006 - ); and TROPOMI onboard the Copernicus S5P (2017 - ). The long-term record (1995 to present) of MINDS datasets makes them very useful for air quality trend studies. Some MINDS datasets with high spatial resolution of only a few kilometers can be used for air quality research and applications at regional scales. In this presentation, we will introduce all of the MINDS products and services, and demonstrate use cases of MINDS data for studying air quality. We will also present a few other NO2 datasets acquired from NASA’s Health and Air Quality Applied Sciences Team (HAQAST), to be archived and distributed by the GES DISC, and highlight some of their applications for air quality and health.

Feng Ding

Enhancing Dataset Discovery With Knowledge Graph Link Prediction Techniques

● In the evolving landscape of open science, the ability to navigate and discover pertinent datasets is increasingly significant. This primarily hinges on the presence of detailed metadata, delineating the dataset’s content, and potential spheres of application. ● The GES DISC datasets are characterized by science keywords to enable dataset discovery in web search interfaces. ● A problem may arise where a dataset lacks a science keyword that it otherwise should have. ● Machine learning techniques such as link prediction can be used to detect these missing science keywords by estimating the probability of new links forming between dataset and keyword nodes.

machine learning

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km x 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

climate projection

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km by 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

general circulation model (GCM)

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES