Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “outlier detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Machine Learning for Automating Analysis of Speckle Dynamics

Code used for generating results in the publication for automating analysis of non-equilibrium X-ray Photon Correlation Spectroscopy data, 'Machine Learning for analysis of speckle dynamics: quantification and outlier detection'.

Konstantinova, Tatiana [Brookhaven National Lab. (↗

PAVC Gridded 20m Alaska NGEE Tier3 PFTs v1.0

These 20-meter spatial resolution gridded products provide per-pixel fractional cover (%) of Next Generation Ecosystem Experiments (NGEE) Arctic Plant Functional Types (PFTs) Tier 3 across Alaska, north of the boreal treeline. The products were developed for the NGEE Arctic project, which is improving Arctic vegetation representation and parameterization of the E3SM Land Model. This dataset includes 8 files containing fractional cover for NGEE Tier 3 PFTs (https://data.ess-dive.lbl.gov/view/doi:10.15485/2529470): (1) bryophytes; (2) lichens; (3) non-vascular plants, i.e., the sum of lichens and bryophytes; (4) deciduous shrubs, (5) evergreen shrubs, (6) forbs, (7) graminoids, and a non-PFT class, (8) litter. Each pixel contains the percent cover (expressed as a fraction of total ground cover) that was predicted by random-forest regression models. The random-forest models were trained on cover data collected at 978 plots from 2010 to 2021, of which are archived in the Pan-Arctic Vegetation Cover (PAVC) database (https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2483557). The plot cover was linked to 20-meter spatial resolution, satellite-derived predictor variables: Sentinel-2 spectra and Sentinel-1 polarizations averaged over the 2019 growing season, as well as topographical features derived from ArcticDEM. Then, spatio-temporally anomalous plot data that introduced large variability to the regression outcomes were dropped using the Cook’s distance outlier detection method, and the models were re-created using high-quality plots and their associated satellite derived explanatory variables per each PFT. The correlations between plot-observed and satellite-derived fractional cover for all PFTs were well correlated (R2 = 0.69–0.95 and 0.5 for litter) and had low RMSE bias (0.02–0.11). This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.

54 ENVIRONMENTAL SCIENCES↗

High-dimensional Data-driven Energy optimization for Multi-Modal Transit Agencies (HD-EMMA) (Final Technical Report)

Public bus transit services in the U.S. are responsible for at least 19.7 million metric tons of CO 2 emission annually. Electric vehicles (EVs) can have a much lower environmental impact than comparable internal combustion engine vehicles (ICEVs), especially in urban areas. Unfortunately, EVs are also much more expensive than ICEVs. As a result, many public transit agencies can afford only mixed fleets of transit vehicles, consisting of EVs, hybrids (HEVs), and ICEVs. Transit agencies that operate such mixed fleets of vehicles face a challenging optimization problem: these agencies need to decide which vehicles are assigned to serving which transit trips. Since the advantage of EVs over ICEVs varies depending on the route and time of day (e.g., the benefit of EVs is higher in slower traffic with frequent stops and lower on highways), the assignment can have a significant effect on energy use and, hence, environmental impact. Through this project, we have developed reference data about energy collections and constructed a set of machine learning models that can accurately predict the energy consumption for the whole fleet at the level of each trip. We have used these models to develop a scheduling and assignment strategy that can rotate the different vehicle types across the transit agencies’ routes. The optimization algorithm ensures that the vehicles are matched to trips considering weather patterns, expected congestion, and road gradients to minimize the overall energy usage. We list the key observations from our project for other practitioners below. Details are available in the report, and the list of source code and our publications are included in the appendix. 1. We have demonstrated the feasibility of collecting, merging and analyzing large volumes of high-resolution real-world telemetry data from a mixed vehicle fleet. To mitigate the inherent noise of the recorded GPS points, the team developed an algorithm that filters data and maps the points onto a street. The algorithm considers previous and subsequent location measurements and different characteristics of nearby streets to determine how likely the vehicle travels on them. Then, the team segmented the time series into disjoint contiguous samples based on adjacent road segments and repeated the outlier detection and removal. For each data point, the team added features corresponding to elevation changes within the samples, weather features, such as temperature, and traffic data, such as speed ratio between actual speed and free-flow speed. 2. We have developed two forms of machine learning models that be used to understand and analyze the energy operations of a mixed vehicle transit fleet. The micro prediction model provides estimates of instantaneous energy prediction for all types of buses (diesel, hybrid, and electric). Such a model is important in evaluating the energy impacts of real-time bus operation strategies, but it is challenging due to diversified driving cycles of transit buses. The model can help the drivers understand the impact of their driving behaviors and short-term congestions. The macro prediction models estimate average energy consumption across the whole trip considering the features: distance traveled, various road-type features, elevation change, day of the week, time of day, various weather features (temperature, humidity, etc.), and traffic features (speed ratio and jam factor). 3. We have demonstrated that it is possible to transfer the machine learning models we have developed in this project to other teams and cities by using inductive transfer learning. We also showed that the performance of the macro energy prediction models can be improved using a multi-task learning approach where the learning parameters are shared between the models being developed for different vehicle types. The advantage of this approach is improved learning performance as the models can exploit common spatio-temporal and environmental characteristics. 4. Finally, we have developed trip and vehicle assignment and scheduling algorithms that use the energy prediction models and develop a trip to vehicle type (diesel, electric, hybrid) assignment for the whole operation to reduce overall emissions and cost. We have shown through simulations that the proposed algorithms can save $\$$ 48,910 in energy costs and 175 metric tons of CO 2 emission annually for CARTA.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

PVAnalytics: A Python Package for Automated Processing of Solar Time Series Data

Multiple publicly available software packages exist that analyze solar time series data, including RdTools and Solar Data Tools, among others. Several of these packages contain their own unique quality assurance (QA) and feature recognition algorithms. The python PVAnalytics package was developed to offer an internally consistent source for these analysis tools, making it easier for the end user to deploy these routines on his or her solar data. The PVAnalytics package currently contains routines for outlier detection, inverter clipping detection, irradiance and temperature checks, orientation checks, and data shift detection, among other functions. These functions have been aggregated from various sources including Solar Forecast Arbiter, RdTools, and the QA process developed by NREL's PV Fleets Initiative. We are continuously adding new functionality to the package, including documentation, examples and algorithms. By bundling QA functionality into a single software package, we hope to make PVAnalytics a comprehensive software library to support analysis of solar metadata and time series data.

data cleaning↗

Genetic monitoring of steelhead in the Klickitat River to estimate productivity, straying, and migration timing

Abstract Objective Salmonids with a complex life history variation present challenges for conservation management, but genetic approaches alongside fisheries monitoring can address questions regarding the viability of the natural populations. Methods We genotyped adult (n = 3108) and juvenile (n = 2624) samples of steelhead Oncorhynchus mykiss that were collected in the Klickitat River, Washington, USA at traps in the lower drainage to examine tributary level productivity, straying from outside sources, and variation in adult migration timing. Result Genetic assignment of steelhead from this system indicated that the majority were produced within or near tributaries of the middle Klickitat River (juvenile mean = 72.8%; adult mean = 87.3%). Analyses with parentage-based tagging identified that most hatchery-origin adults were assigned to the Skamania Hatchery (80.8%), as expected, since this has been the release stock for decades within the Klickitat River drainage. Hatchery-origin adults were also identified from programs operating outside the Klickitat River, which were primarily strays from Snake River hatcheries. Most natural-origin steelhead were assigned to the Klickitat River, but there were also natural-origin fish identified as strays from other regions of the Columbia River (22.3% of natural returns). We also examined genes known to be associated with migration timing in adult steelhead observed at the trap and observed a strong relationship between migration date and alleles for early and late migration, but individual outliers were detected across seasons. Conclusion Our results indicate that genetic variation of steelhead in the Klickitat River has been influenced by hatchery programs as well as natural-origin straying from other subbasins, but genetic diversity remains high throughout the subbasin, and both early and late migration alleles are maintained. The genetic diversity present in Klickitat River steelhead may enable this Endangered Species Act listed (threatened) species to better adapt to stochastic environmental conditions compared to less diverse populations.

Collins, Erin E. (ORCID:0000000260978479)↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

A Probabilistic Autoencoder for Type Ia Supernova Spectral Time Series

We construct a physically parameterized probabilistic autoencoder (PAE) to learn the intrinsic diversity of Type Ia supernovae (SNe Ia) from a sparse set of spectral time series. The PAE is a two-stage generative model, composed of an autoencoder that is interpreted probabilistically after training using a normalizing flow. We demonstrate that the PAE learns a low-dimensional latent space that captures the nonlinear range of features that exists within the population and can accurately model the spectral evolution of SNe Ia across the full range of wavelength and observation times directly from the data. By introducing a correlation penalty term and multistage training setup alongside our physically parameterized network, we show that intrinsic and extrinsic modes of variability can be separated during training, removing the need for the additional models to perform magnitude standardization. We then use our PAE in a number of downstream tasks on SNe Ia for increasingly precise cosmological analyses, including the automatic detection of SN outliers, the generation of samples consistent with the data distribution, and solving the inverse problem in the presence of noisy and incomplete data to constrain cosmological distance measurements. We find that the optimal number of intrinsic model parameters appears to be three, in line with previous studies, and show that we can standardize our test sample of SNe Ia with an rms of 0.091 ± 0.010 mag, which corresponds to 0.074 ± 0.010 mag if peculiar velocity contributions are removed.

79 ASTRONOMY AND ASTROPHYSICS↗

Searching for Novel Chemistry in Exoplanetary Atmospheres Using Machine Learning for Anomaly Detection

Abstract The next generation of telescopes will yield a substantial increase in the availability of high-quality spectroscopic data for thousands of exoplanets. The sheer volume of data and number of planets to be analyzed greatly motivate the development of new, fast, and efficient methods for flagging interesting planets for reobservation and detailed analysis. We advocate the application of machine learning (ML) techniques for anomaly (novelty) detection to exoplanet transit spectra, with the goal of identifying planets with unusual chemical composition and even searching for unknown biosignatures. We successfully demonstrate the feasibility of two popular anomaly detection methods (local outlier factor and one-class support vector machine) on a large public database of synthetic spectra. We consider several test cases, each with different levels of instrumental noise. In each case, we use receiver operating characteristic curves to quantify and compare the performance of the two ML techniques.

Astronomy & Astrophysics↗

Evolutionary Analyses of Gene Expression Divergence in Panicum hallii : Exploring Constitutive and Plastic Responses Using Reciprocal Transplants

Abstract The evolution of gene expression is thought to be an important mechanism of local adaptation and ecological speciation. Gene expression divergence occurs through the evolution of cis- polymorphisms and through more widespread effects driven by trans-regulatory factors. Here, we explore expression and sequence divergence in a large sample of Panicum hallii accessions encompassing the species range using a reciprocal transplantation experiment. We observed widespread genotype and transplant site drivers of expression divergence, with a limited number of genes exhibiting genotype-by-site interactions. We used a modified FST–QST outlier approach (QPC analysis) to detect local adaptation. We identified 514 genes with constitutive expression divergence above and beyond the levels expected under neutral processes. However, no plastic expression responses met our multiple testing correction as QPC outliers. Constitutive QPC outlier genes were involved in a number of developmental processes and responses to abiotic environments. Leveraging earlier expression quantitative trait loci results, we found a strong enrichment of expression divergence, including for QPC outliers, in genes previously identified with cis and cis–environment interactions but found no patterns related to trans-factors. Population genetic analyses detected elevated sequence divergence of promoters and coding sequence of constitutive expression outliers but little evidence for positive selection on these proteins. Our results are consistent with a hypothesis of cis-regulatory divergence as a primary driver of expression divergence in P. hallii.

3′ TagSeq↗

Signatures of Selection for Resistance/Tolerance to Perkinsus olseni in Grooved Carpet Shell Clam ( Ruditapes decussatus ) Using a Population Genomics Approach

ABSTRACT The grooved carpet shell clam ( Ruditapes decussatus ) is a bivalve of high commercial value distributed throughout the European coast. Its production has suffered a decline caused by different factors, especially by the parasite Perkinsus olsenii . Improving production of R . decussatus requires genomic resources to ascertain the genetic factors underlying resistance/tolerance to P. olseni i . In this study, the first reference genome of R . decussatus was assembled through long‐ and short‐read sequencing (1677 contigs; 1.386 Mb) and further scaffolded at chromosome level with Hi‐C (19 superscaffolds; 95.4% of assembly). Repetitive elements were identified (32%) and masked for annotation of 38,276 coding‐ and 13,056 non‐coding genes. This genome was used as a reference to develop a 2bRAD‐Seq 13,438 SNP panel for a genomic screening on six shellfish beds distributed across the Atlantic Ocean and Mediterranean Sea. Beds were selected by perkinsosis prevalence and the infection level was individually evaluated in all the samples. Genetic diversity was significantly higher in the Mediterranean than in the Atlantic region. The main genetic breakage was detected between those regions (F ST = 0.224), being the Mediterranean more heterogeneous than the Atlantic. Several loci under divergent selection (394 outliers; 261 genomic windows) were detected across shellfish beds. Samples were also inspected to detect signals of selection for resistance/tolerance to P. olseni i by using infection‐level and population‐genomics approaches, and 90 common divergent outliers for resistance/tolerance to perkinsosis were identified and used for gene mining. Candidate genes and markers identified provide invaluable information for controlling perkinsosis and for improving production of the grooved carpet shell clam.

Sambade, Inés M. [Department of Zoology, Genetics ↗

Method for automatic correction of offset drift in online sensors

Abstract Successful operation and optimization of water treatment systems hinge on the availability of high-quality online sensor measurements. Ideally, the available measurements should be simultaneously accurate (i.e., unbiased and precise), representative, voluminous, and timely. This remains a pain-point in current water infrastructures, forming a barrier to a wider adoption of advanced and autonomous control systems. While short-lived symptoms, such as outliers and spikes, can be detected or corrected with state-of-the-art tools for fault detection and identification, it is much more difficult to detect, diagnose, and correct the symptoms of slow faults, such as changes in offset or sensitivity due to drift. The time scale of drift is often longer than the time scales of the system dynamics of interest. Moreover, sensor drift has been shown to occur at the same time and with similar rates when sensors are exposed to the same conditions. This challenges data quality management strategies based on redundancy. In this contribution, we develop a new method, including both a hands-off sensor calibration mechanism and an information-seeking control architecture that can handle the unique challenge of simultaneous and similar drift in online sensors.

Chowdhury, Dhrubajit↗

A three-year dataset supporting research on building energy management and occupancy analytics

Abstract This paper presents the curation of a monitored dataset from an office building constructed in 2015 in Berkeley, California. The dataset includes whole-building and end-use energy consumption, HVAC system operating conditions, indoor and outdoor environmental parameters, as well as occupant counts. The data were collected during a period of three years from more than 300 sensors and meters on two office floors (each 2,325 m 2 ) of the building. A three-step data curation strategy is applied to transform the raw data into research-grade data: (1) cleaning the raw data to detect and adjust the outlier values and fill the data gaps; (2) creating the metadata model of the building systems and data points using the Brick schema; and (3) representing the metadata of the dataset using a semantic JSON schema. This dataset can be used in various applications—building energy benchmarking, load shape analysis, energy prediction, occupancy prediction and analytics, and HVAC controls—to improve the understanding and efficiency of building operations for reducing energy use, energy costs, and carbon emissions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Where’s Swimmy?: Mining unique color features buried in galaxies by deep anomaly detection using Subaru Hyper Suprime-Cam data

Abstract We present the Swimmy (Subaru WIde-field Machine-learning anoMalY) survey program, a deep-learning-based search for unique sources using multicolored (grizy) imaging data from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). This program aims to detect unexpected, novel, and rare populations and phenomena, by utilizing the deep imaging data acquired from the wide-field coverage of the HSC-SSP. This article, as the first paper in the Swimmy series, describes an anomaly detection technique to select unique populations as “outliers” from the data-set. The model was tested with known extreme emission-line galaxies (XELGs) and quasars, which consequently confirmed that the proposed method successfully selected $\sim\!\! 60\%$–$70\%$ of the quasars and $60\%$ of the XELGs without labeled training data. In reference to the spectral information of local galaxies at z = 0.05–0.2 obtained from the Sloan Digital Sky Survey, we investigated the physical properties of the selected anomalies and compared them based on the significance of their outlier values. The results revealed that XELGs constitute notable fractions of the most anomalous galaxies, and certain galaxies manifest unique morphological features. In summary, deep anomaly detection is an effective tool that can search rare objects, and, ultimately, unknown unknowns with large data-sets. Further development of the proposed model and selection process can promote the practical applications required to achieve specific scientific goals.

Astronomy & Astrophysics↗

Stragglers of the thick disc

Young alpha-rich (YAR) stars have been detected in the past as outliers to the local age - [α/Fe] relation. These objects are enhanced in α-elements, but they are apparently younger than typical thick disc stars. Here, we study the global kinematics and chemical properties of YAR giant stars in the APOGEE DR17 survey and show that they have properties similar to those of the standard thick disc stellar population. This leads us to conclude that YAR are rejuvenated thick disc objects, and the most likely explanation is that they are evolved blue stragglers. This is confirmed by their position in the Hertzsprung–Russel diagram (HRD). Extending our selection to dwarfs allowed us to obtain the first general straggler distribution in an HRD of field stars. We also compared the elemental abundances of our sample with those of standard thick disc stars and found that our YAR stars are shifted in oxygen, magnesium, sodium, and the slow neutron-capture element cerium. Although we detected no sign of binarity for most objects, the enhancement in cerium may be a signature of a mass transfer from an asymptotic giant branch companion. The most massive YAR stars suggest that mass transfer from an evolved star may not be the only plausible formation pathway and that other scenarios, such as collision or coalescence, should be considered.

79 ASTRONOMY AND ASTROPHYSICS↗

Challenging problems of quality assurance and quality control (QA/QC) of meteorological time series data

Abstract Representativeness and quality of collected meteorological data impact accuracy and precision of climate, hydrological, and biogeochemical analyses and predictions. We developed a comprehensive Quality Assurance (QA) and Quality Control (QC) statistical framework, consisting of three major phases: Phase I—Preliminary data exploration, i.e., processing of raw datasets, with the challenging problems of time formatting and combining datasets of different lengths and different time intervals; Phase II—QA of the datasets, including detecting and flagging of duplicates, outliers, and extreme data; and Phase III—the development of time series of a desired frequency, imputation of missing values, visualization and a final statistical summary. The paper includes two use cases based on the time series data collected at the Billy Barr meteorological station (East River Watershed, Colorado), and the Barro Colorado Island (BCI, Panama) meteorological station. The developed statistical framework is suitable for both real-time and post-data-collection QA/QC analysis of meteorological datasets.

54 ENVIRONMENTAL SCIENCES↗