Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “outlier detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

On The Processing of Log Files for Monitoring Antenna Health

In order to improve the quality of geodetic results, we have developed an infrastructure for timely processing of telemetry from IVS observing stations. We check every hour for new log files with telemetry from both VLBI observing sessions, single dish experiments, and stow-in data collection and automatically process them. The telemetry data we use is the system temperature, phase calibration phases and amplitudes, system equivalent flux density, and the differences between formatter clock and GPS clock. For the system temperature and phase calibration, processing includes filtering out outliers and computing averages and rms of the scatter in each scan. Furthermore, for the phase calibration we also compute the group delay and detect spurious signals. Cleaned and post-processed telemetry is archived. Our process detects abnormalities, such as, anomalously high system temperature, unstable phase calibration phases, jumps in the GPS and formatter clock differences, and others. With our procedure, the latency of detection of station abnormalities is reduced to less than two hours. Early detection of abnormalities reduces the amount of affected data since station personnel get early alerts. We discuss our experience of running this system since 2022.

Phase Calibration↗

On The Processing of Log Files for Monitoring Antenna Health

In order to improve the quality of geodetic results, we have developed an infrastructure for timely processing of telemetry from IVS observing stations. We check every hour for new log files with telemetry from both VLBI observing sessions, single dish experiments, and stow-in data collection and automatically process them. The telemetry data we use is the system temperature, phase calibration phases and amplitudes, system equivalent flux density, and the differences between formatter clock and GPS clock. For the system temperature and phase calibration, processing includes filtering out outliers and computing averages and rms of the scatter in each scan. Furthermore, for the phase calibration we also compute the group delay and detect spurious signals. Cleaned and post-processed telemetry is archived. Our process detects abnormalities, such as, anomalously high system temperature, unstable phase calibration phases, jumps in the GPS and formatter clock differences, and others. With our procedure, the latency of detection of station abnormalities is reduced to less than two hours. Early detection of abnormalities reduces the amount of affected data since station personnel get early alerts. We discuss our experience of running this system since 2022.

VLBI↗

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗

Anomaly detection in the Zwicky Transient Facility DR3

We present results from applying the SNAD anomaly detection pipeline to the third public data release of the Zwicky Transient Facility (ZTF DR3). The pipeline is composed of three stages: feature extraction, search of outliers with machine learning algorithms, and anomaly identification with followup by human experts. Our analysis concentrates in three ZTF fields, comprising more than 2.25 million objects. A set of four automatic learning algorithms was used to identify 277 outliers, which were subsequently scrutinized by an expert. From these, 188 (68 per cent) were found to be bogus light curves – including effects from the image subtraction pipeline as well as overlapping between a star and a known asteroid, 66 (24 per cent) were previously reported sources whereas 23 (8 per cent) correspond to non-catalogued objects, with the two latter cases of potential scientific interest (e.g. one spectroscopically confirmed RS Canum Venaticorum star, four supernovae candidates, one red dwarf flare). Moreover, using results from the expert analysis, we were able to identify a simple bi-dimensional relation that can be used to aid filtering potentially bogus light curves in future studies. We provide a complete list of objects with potential scientific application so they can be further scrutinised by the community. These results confirm the importance of combining automatic machine learning algorithms with domain knowledge in the construction of recommendation systems for astronomy. Our code is publicly available.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantifying the Impact of LSST u -band Survey Strategy on Photometric Redshift Estimation and the Detection of Lyman-break Galaxies

The Vera C. Rubin Observatory will conduct the Legacy Survey of Space and Time (LSST), promising to discover billions of galaxies out to redshift 7, using six photometric bands (ugrizy) spanning the near-ultraviolet to the near-infrared. The exact number of and quality of information about these galaxies will depend on survey depth in these six bands, which in turn depends on the LSST survey strategy, i.e., how often and how long to expose in each band. u-band depth is especially important for photometric redshift (photo-z) estimation and for detection of high-redshift Lyman-break galaxies (LBGs). In this paper, we use a simulated galaxy catalog and an analytic model for the LBG population to study how recent updates and proposed changes to Rubin’s u-band throughput and LSST survey strategy impact photo-z accuracy and LBG detection. We find that proposed variations in u-band strategy have a small impact on photo-z accuracy for z < 1.5 galaxies, but the outlier fraction, scatter, and bias for higher-redshift galaxies vary by up to 50%, depending on the survey strategy considered. The number of u-band dropout LBGs at z ∼ 3 is also highly sensitive to the u-band depth, varying by up to 500%, while the number of griz-band dropouts is only modestly affected. Under the new u-band strategy recommended by the Rubin Survey Cadence Optimization Committee, we predict u-band dropout number densities of 110 deg −2 (3200 deg −2 ) in year 1 (10) of LSST. We discuss the implications of these results for LSST cosmology.

79 ASTRONOMY AND ASTROPHYSICS↗

Regional and Temporal Variability of Atmospheric River Seasonality: Influences of Detection Algorithms and Moisture Transport Dynamics

Abstract Understanding the regional and temporal variability of atmospheric river (AR) seasonality is crucial for preparedness and mitigation of extreme events. While ARs were thought to peak in winter, recent research shows they exhibit region‐specific seasonality and are heavily influenced by the chosen detection algorithm. This study examines the link between the year‐to‐year consistency of peak‐AR activity to the presence of a dominant seasonal pattern, considering both location and algorithm choice. Regions are categorized by their temporal characteristics: consistent patterns (e.g., East Asia), patterns with occasional outliers (e.g., British Columbia coast), and regions lacking a clear dominant peak season (e.g., South Atlantic, parts of Australia). Hence, not all regions display a consistent seasonal cycle of AR activity. This study quantifies the extent to which a region experiences a dominant peak season of AR activity (or lacks one) and offers insights to enhance decision‐making in water management, natural hazard preparedness, and forecasting. Furthermore, given our finding that detection algorithms influence the peak season of AR activity, we also examine two diagnostic variables representative of moisture transport to corroborate our results. Integrated vapor transport, which captures meridional and zonal moisture transport, and Moist Wave Activity, representing moisture intrusions from lower to higher latitudes, are examined. Our analysis indicates that inconsistencies in the seasonal cycle of AR activity are not solely due to discrepancies in detection algorithms but also arise from changes in moisture transport. Plain Language Summary Atmospheric rivers (ARs) are critical weather phenomena that can cause extreme events like heavy rainfall and flooding. Understanding when and where ARs are most likely to occur throughout the year is essential for preparing and responding to these events. Traditionally, ARs were thought to peak in winter, but recent studies show this varies by region. Our study helps address the challenge decision‐makers face in anticipating and preparing for AR events by providing insights into the consistency of peak seasonal patterns across different areas. Some regions, like East Asia, and the British Columbia coast, show a consistent peak season, while others, like the South Atlantic and parts of Australia, have significant year‐to‐year variations, making it hard to identify a dominant season. To better understand these changes over time, the study also examines how moisture moves in the atmosphere, using Integrated Vapor Transport (which looks at moisture movement in various directions) and Moist Wave Activity (which tracks moisture shifts from lower to higher latitudes). The findings suggest that inconsistencies in AR patterns are due not only to detection methods but also due to changes in moisture transport. Key Points The peak season of atmospheric river activity can change depending on the year in some areas Interannual variations in the peak season can make identifying a dominant season challenging for some regions Frequent shifts in peak season across years reflect inconsistencies tied to detection algorithms and to underlying dynamics

Kamnani, Diya↗

Automatic Drift Correction through Nonlinear Sensing

For successful design and operation of advanced monitoring and control systems, engineers rely on high quality sensor signals that are simultaneously accurate, representative, voluminous, and timely. Unfortunately, sensor faults are common and lead to short-lived symptoms, such as outliers and spikes as well as long-lived symptoms, such as sensor drift. Sensor drift belongs to the category of incipient faults. These are particularly challenging to detect, diagnose, and correct as the time scales of these faults are typically longer than the time scales of the system dynamics that are of interest. Moreover, if sensor drift occurs as a result of exposure to measured medium, then it is likely that multiple sensors will exhibit similar drift rates, thus challenging fault management strategies based on redundancy. In this contribution, we present a first method that can handle this unique challenge.

Chowdhury, Dhruba↗

Know Your Space: Inlier and Outlier Construction for Calibrating Medical OOD Detectors

This software offers methods and functions for training calibrated out-of-distribution detectors for medical image classification tasks. It includes functionalities for training, synthesizing data augmentations, calibration, and out-of-distribution detection. Developed using PyTorch, this software is compatible with standard neural network architectures used for imaging data. Additionally, it provides capabilities to compute evaluation metrics for assessing the performance and quality of the detectors.

Narayanaswamy, VivekSivaraman↗

Multiple-Beam Detection of Fast Transient Radio Sources

A method has been designed for using multiple independent stations to discriminate fast transient radio sources from local anomalies, such as antenna noise or radio frequency interference (RFI). This can improve the sensitivity of incoherent detection for geographically separated stations such as the very long baseline array (VLBA), the future square kilometer array (SKA), or any other coincident observations by multiple separated receivers. The transients are short, broadband pulses of radio energy, often just a few milliseconds long, emitted by a variety of exotic astronomical phenomena. They generally represent rare, high-energy events making them of great scientific value. For RFI-robust adaptive detection of transients, using multiple stations, a family of algorithms has been developed. The technique exploits the fact that the separated stations constitute statistically independent samples of the target. This can be used to adaptively ignore RFI events for superior sensitivity. If the antenna signals are independent and identically distributed (IID), then RFI events are simply outlier data points that can be removed through robust estimation such as a trimmed or Winsorized estimator. The alternative "trimmed" estimator is considered, which excises the strongest n signals from the list of short-beamed intensities. Because local RFI is independent at each antenna, this interference is unlikely to occur at many antennas on the same step. Trimming the strongest signals provides robustness to RFI that can theoretically outperform even the detection performance of the same number of antennas at a single site. This algorithm requires sorting the signals at each time step and dispersion measure, an operation that is computationally tractable for existing array sizes. An alternative uses the various stations to form an ensemble estimate of the conditional density function (CDF) evaluated at each time step. Both methods outperform standard detection strategies on a test sequence of VLBA data, and both are efficient enough for deployment in real-time, online transient detection applications.

Thompson, David R.↗

An Analysis of the Statistics and Systematics of Limb Anomaly Detections in HST/STIS Transit Images of Europa

Several recent studies derived the existence of plumes on Jupiter’s moon Europa. The only technique that provided multiple detections is the far-ultraviolet imaging observations of Europa in transit of Jupiter taken by the Space Telescope Imaging Spectrograph (STIS) on the Hubble Space Telescope (HST). In this study, we reanalyze the three HST/STIS transit images in which Sparks et al. identified limb anomalies as evidence for Europa’s plume activity. After reproducing the results of Sparks et al., we find that positive outliers are similarly present in the images as the negative outliers that were attributed to plume absorption. A physical explanation for the positive outliers is missing. We then investigate the systematic uncertainties and statistics in the images and identify two factors that are crucial when searching for anomalies around the limb. One factor is the alignment between the actual and assumed locations of Europa on the detector. A misalignment introduces distorted statistics, most strongly affecting the limb above the darker trailing hemisphere where the plumes were detected. The second factor is a discrepancy between the observation and the model used for comparison, adding uncertainty in the statistics. When accounting for these two factors, the limb minima (and maxima) are consistent with random statistical occurrence in a sample size given by the number of pixels in the analyzed limb region. The plume candidate features in the three analyzed images can be explained by purely statistical fluctuations and do not provide evidence for absorption by plumes.

79 ASTRONOMY AND ASTROPHYSICS↗

Algorithm for Identifying Erroneous Rain-Gauge Readings

An algorithm analyzes rain-gauge data to identify statistical outliers that could be deemed to be erroneous readings. Heretofore, analyses of this type have been performed in burdensome manual procedures that have involved subjective judgements. Sometimes, the analyses have included computational assistance for detecting values falling outside of arbitrary limits. The analyses have been performed without statistically valid knowledge of the spatial and temporal variations of precipitation within rain events. In contrast, the present algorithm makes it possible to automate such an analysis, makes the analysis objective, takes account of the spatial distribution of rain gauges in conjunction with the statistical nature of spatial variations in rainfall readings, and minimizes the use of arbitrary criteria. The algorithm implements an iterative process that involves nonparametric statistics.

Rickman, Doug↗

Are Soft Short Tests Good Indicators of Internal Li-ion Cell Defects?

The self discharge test at full state of charge, may not be a good one to detect subtle defects since the li-ion chemistry has the highest self discharge at full state of charge. One should characterize self discharge versus storage time for each cell manufacturer/design to differentiate between normal self discharge and that due to a subtle manufacturing defect. The various soft short test methods indicate that if this test is carried out at full discharge (0% SOC) with all capacity removed (by lowering the current load in a stepwise manner to the same end of discharge voltage), then the cells need to be placed in storage for more than 72 hours to get a good analysis on the presence of subtle defects since it takes more than 72 hours to achieve voltage stabilization. If the cells are to be charged up even to a small percentage (ex. 1%), 72 hours are sufficient to determine issues. However, the pass/fail criteria should be based on a valid OCV decline. Less than 10 mV voltage decline is not a good method to detect subtle defects. As mentioned in the first bullet, self discharge is a competing reaction when a charge is introduced and hence a characterization of the self discharge versus storage time is required to fully correlate voltage decline to a failure due to a subtle defect. Soft short test method cannot be relied on for defect detection because cells with and without voltage decline seemed to have similar defects and characteristics. Screening methods such as internal resistance and capacity as well as a 3-sigma range for OCV, mass and dimensions should be used to screen out outliers. A very critical aspect in the understanding of subtle defects is to carry out destructive analysis of cells from every lot to confirm the quality of production and screen all cells and batteries in a stringent manner to have a high quality set of flight cells. Self Discharge Test: Fully charged cells shall be placed in Open circuit stand for 72 hours (OCV measurement twice a day); continue for total of 14 days with 1 reading per day 2. Soft Short Test 1: Fully charge; cells discharged to manufacturer's end of disch. Voltage (EODV) cutoff at C/5 rate; stand for 30 minutes; discharge with C/500 to the same EODV. stand for another 30 minutes; discharge the cells again using C/1000 current to the same EODV. OCV measurements twice a day for 72 hours and then for total of 14 days (data collection same as in 1.) 3. Soft Short Test 2: Fully charge; cells discharged to the manuf. EODV with a C/13 constant current; provide a 10 hour rest, discharge again to the same EODV with a current of C/250, provide a 10 hour rest, discharge again using a C/250 rate, provide a 24 hour rest, charge using C/250 to 3.15 V (for ~12 hours). OCV measurements twice a day for at 72 hours. (data collection same as in 1.) 4. Soft Short Test 3: Fully charge; cells shall be discharged using C/10 current to manuf. EODV. Allow the cell to remain at Open circuit for 10 seconds. Discharge the cell at C/20 rate to the same EODV, hold open circuit for 24 hours. Discharge the cells at C/200 rate to the same end of voltage cutoff and hold open circuit for 24 hours. Discharge the cells one more time at C/200 rate to the same EODV and hold open circuit for 36 hours. Charge at C/200 rate to 3.15 V and hold for 3 days. Record OCV during the open circuit stand periods every 12 hours and at the beginning and end of the 3 day hold (include the 12 hour OCV recording during this time also). Capacity Cycling: Cells with declining voltages - one cell from each manufacturer chosen for cycling Destructive Physical Analysis (DPA): Cells with and without decline chosen from each lot for DPA.

Jeevarajan, J.↗

Photo-z outlier self-calibration in weak lensing surveys

Calibrating photometric redshift errors in weak lensing surveys with external data is extremely challenging. In this work, we show that both Gaussian and outlier photo-z parameters can be self-calibrated from the data alone. This comes at no cost for the neutrino masses, curvature and dark energy equation of state w0, but with a 65% degradation when both w0 and wa are varied. We perform a realistic forecast for the Vera Rubin Observatory (VRO) Legacy Survey of Space and Time (LSST) 3× 2 analysis, combining cosmic shear, projected galaxy clustering and galaxy - galaxy lensing. We confirm the importance of marginalizing over photo-z outliers. We examine a subset of internal cross-correlations, dubbed "null correlations", which are usually ignored in 3× 2 analyses. Despite contributing only ~ 10% of the total signal-to-noise, these null correlations improve the constraints on photo-z parameters by up to an order of magnitude. Using the same galaxy sample as sources and lenses dramatically improves the photo-z uncertainties too. Together, these methods add robustness to any claim of detected new Physics, and reduce the statistical errors on cosmology by 15% and 10% respectively. Finally, including CMB lensing from an experiment like Simons Observatory or CMB-S4 improves the cosmological and photo-z posterior constraints by about 10%, and further improves the robustness to systematics. To give intuition on the Fisher forecasts, we examine in detail several toy models that explain the origin of the photo-z self-calibration. Our Fisher code LaSSI (Large-Scale Structure Information), which includes the effect of Gaussian and outlier photo-z, shear multiplicative bias, linear galaxy bias, and extensions to LCDM, is publicly available at \href{https://github.com/EmmanuelSchaan/LaSSI}{https://github.com/EmmanuelSchaan/LaSSI}.

79 ASTRONOMY AND ASTROPHYSICS↗

QSO photometric redshifts using machine learning and neural networks

ABSTRACT The scientific value of the next generation of large continuum surveys would be greatly increased if the redshifts of the newly detected sources could be rapidly and reliably estimated. Given the observational expense of obtaining spectroscopic redshifts for the large number of new detections expected, there has been substantial recent work on using machine learning techniques to obtain photometric redshifts. Here, we compare the accuracy of the predicted photometric redshifts obtained from deep learning (DL) with the k-nearest neighbour (kNN) and the decision tree regression (DTR) algorithms. We find using a combination of near-infrared, visible, and ultraviolet magnitudes, trained upon a sample of Sloan Digital Sky Survey quasi-stellar objects, that the kNN and DL algorithms produce the best self-validation result with a standard deviation of σΔz = 0.24 (σΔz(norm) = 0.11). Testing on various subsamples, we find that the DL algorithm generally has lower values of σΔz, in addition to exhibiting a better performance in other measures. Our DL method, which uses an easy to implement off-the-shelf algorithm with neither filtering nor removal of outliers, performs similarly to other, more complex, algorithms, resulting in an accuracy of Δz < 0.1 up to z ∼ 2.5. Applying the DL algorithm trained on our 70 000 strong sample to other independent (radio-selected) data sets, we find σΔz ≤ 0.36 (σΔz(norm) ≤ 0.17) over a wide range of radio flux densities. This indicates much potential in using this method to determine photometric redshifts of quasars detected with the Square Kilometre Array.

Curran, S. J.↗

Redshifts of radio sources in the Million Quasars Catalogue from machine learning

ABSTRACT With the aim of using machine learning techniques to obtain photometric redshifts based upon a source’s radio spectrum alone, we have extracted the radio sources from the Million Quasars Catalogue. Of these, 44 119 have a spectroscopic redshift, required for model validation, and for which photometry could be obtained. Using the radio spectral properties as features, we fail to find a model which can reliably predict the redshifts, although there is the suggestion that the models improve with the size of the training sample. Using the near-infrared–optical–ultraviolet bands magnitudes, we obtain reliable predictions based on the 12 503 radio sources which have all of the required photometry. From the 80:20 training–validation split, this gives only 2501 validation sources, although training the sample upon our previous SDSS model gives comparable results for all 12 503 sources. This makes us confident that SkyMapper, which will survey southern sky in the u, v, g, r, i, z bands, can be used to predict the redshifts of radio sources detected with the Square Kilometre Array. By using machine learning to impute the magnitudes missing from much of the sample, we can predict the redshifts for 32 698 sources, an increase from 28 to 74 per cent of the sample, at the cost of increasing the outlier fraction by a factor of 1.4. While the ‘optical’ band data prove successful, at this stage we cannot rule out the possibility of a radio photometric redshift, given sufficient data which may be necessary to overcome the relatively featureless radio spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

JIRIAF: JLAB Integrated Research Infrastructure Acros Facilities

The JIRIAF project aims to combine geographically diverse computing facilities into an integrated science infrastructure. This project starts by dynamically evaluating temporarily unallocated or idled compute resources from multiple providers. These resources are integrated to handle additional workloads without affecting local running jobs. This paper describes our approach to launch best-effort batch tasks which exploit these underutilized resources. Our system measures the real-time behavior of jobs running on a machine and learns to distinguish typical performance from outliers. Unsupervised ML techniques are used to analyze hardware-level performance measures, followed by a real-time cross-correlation analysis to determine which applications cause performance degradation. We then ameliorate bad behavior by throttling these processes. We demonstrate that problematic performance interference can be detected and acted on, which makes it possible to continue to share resources between applications and simultaneously maintain high utilization levels in a computing cluster. We relocate the CLAS12 data processing workflow to a remote data center for a case study, preventing file migration and temporal data persistency.

Lawrence, David↗

Updating the Standard Spatial Observer for Contrast Detection

Watson and Ahmuada (2005) constructed a Standard Spatial Observer (SSO) model for foveal luminance contrast signal detection based on the Medelfest data (Watson, 1999). Here we propose two changes to the model, dropping the oblique effect from the CSF and using the cone density data of Curcio et al. (1990) to estimate the variation of sensitivity with eccentricity. Dropping the complex images, and using medians to exclude outlier data points, the SSO model now accounts for essentially all the predictable variance in the data, with an RMS prediction error of only 0.67 dB.

Ahumada, Albert J.↗

Anomaly Detection for the Roman Space Telescope Wide Field Instrument’s Science Data Processing Pipeline

The Roman Space Telescope (RST) Wide Field Instrument (WFI) will be utilizing a preliminary Science Data Processing (SDP) pipeline during its Integration and Test, and to some extent during Operations, to track basic statistics and identify known features such as cosmic rays, snowballs as well as possible anomalies in raw detector data. In our detectors, these anomalies appear as jumps in the ramp of a readout and are classified as cosmic rays if they appear as a streak or snowballs if they’re more circular. The WFI employs an array of 18 H4RG-10 detectors that collect image samples. Each set of raw frames within a non-destructive exposure is packaged by the SDP pipeline into image cubes for each detector. Each cube is a time series of 4096 × 4096 accumulating pixel frames. The preliminary analysis pipeline is used to locate anomalies in these time-series accumulation frames and identify the type of anomaly, either natural phenomena or detector characteristic. To compare different methods, we’ve implemented both heuristic-based and data-driven methods to identify anomalies. For the heuristic-based approach, we identify snowballs and cosmic rays by the size and shape of outlier pixel clusters between consecutive frames. For data driven methods, we evaluated a Convolutional Neural Network (CNN) model, and more traditional methods like Principal Component Analysis (PCA). CNN is a supervised learning/classification method. Thus, we used a labeled dataset of anomalies to perform segmentation of the image and identify anomalies. We used previously identified cosmic rays and snowballs to measure the accuracy and efficiency of the mentioned approaches. In evaluating these methods, we aim to pick the best fit for the SDP pipeline’s anomaly detection in terms of both performance and runtime.

Paul Horton↗