Engineering PapersSearch

SEARCH · Engineering Papers

Results for “ARIMA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Structure of the Melajo clay near Arima, Trinidad and strike-slip motion in the El Pilar fault zone

No consensus has yet emerged on the sense, timing and amount of motion in the El Pilar fault zone. As a contribution to the study of this problem, a critical area within the zone in North Central Trinidad has been mapped. On the basis of the mapping, it is concluded that the El Pilar zone has been active in right-lateral strike-slip motion during the Pleistocene. Recognition of structural styles akin to those of the mapped area leads to the suggestion that the El Pilar zone is part of a 300 km wide plate boundary zone extending from the Orinoco delta northward to Grenada. Lateral motion of the Caribbean plate with respect to South America has been suggested to amount to 1900 km in the last 38 Ma. Part of this displacement since the Miocene can be readily accommodated within the broad zone identified here. No one fault system need account for more than a fraction of the total motion and all faults need not be active simultaneously.

Robertson, P.

Temporal Forecasting of Distributed Temperature Sensing in a Thermal Hydraulic System With Machine Learning and Statistical Models

We benchmark performance of long-short term memory (LSTM) network machine learning model and autoregressive integrated moving average (ARIMA) statistical model in temporal forecasting of distributed temperature sensing (DTS). Data in this study consists of fluid temperature transient measured with two co-located Rayleigh scattering fiber optic sensors (FOS) in a forced convection mixing zone of a thermal tee. We treat each gauge of a FOS as an independent temperature sensor. We first study prediction of DTS time series using Vanilla LSTM and ARIMA models trained on prior history of the same FOS that is used for testing. The results yield maximum absolute percentage error (MaxAPE) and root mean squared percentage error (RMSPE) of 1.58% and 0.06% for ARIMA, and 3.14% and 0.44% for LSTM, respectively. Next, we investigate zero-shot forecasting (ZSF) with LSTM and ARIMA trained on history of the co-located FOS only, which is advantageous when limited training data is available. The ZSF MaxAPE and RMSPE values for ARIMA are comparable to those of the Vanilla use case, while the error values for LSTM increase. We show that in ZSF, performance of LSTM network can be improved by training on most correlated gauges between the two FOS, which are identified by calculating the Pearson correlation coefficient. The improved ZSF MaxAPE and RMSPE for LSTM are 4.4% and 0.33%, respectively. Performance of ZSF LSTM can be further enhanced through transfer learning (TL), where LSTM is re-trained on a subset of the FOS that is the target of forecasting. We show that LSTM pre-trained on correlated dataset and re-trained on 30% of testing target dataset achieves MaxAPE and RMSPE values of 2.32% and 0.28%, respectively.

ARIMA

Comparison of Machine Learning-Based Predictive Models of the Nutrient Loads Delivered from the Mississippi/Atchafalaya River Basin to the Gulf of Mexico

Predicting nutrient loads is essential to understanding and managing one of the environmental issues faced by the northern Gulf of Mexico hypoxic zone, which poses a severe threat to the Gulf’s healthy ecosystem and economy. The development of hypoxia in the Gulf of Mexico is strongly associated with the eutrophication process initiated by excessive nutrient loads. Due to the complexities in the excessive nutrient loads to the Gulf of Mexico, it is challenging to understand and predict the underlying temporal variation of nutrient loads. The study was aimed at identifying an optimal predictive machine learning model to capture and predict nonlinear behavior of the nutrient loads delivered from the Mississippi/Atchafalaya River Basin (MARB) to the Gulf of Mexico. For this purpose, monthly nutrient loads (N and P) in tons were collected from US Geological Survey (USGS) monitoring station 07373420 from 1980 to 2020. Machine learning models—including autoregressive integrated moving average (ARIMA), gaussian process regression (GPR), single-layer multilayer perceptron (MLP), and a long short-term memory (LSTM) with the single hidden layer—were developed to predict the monthly nutrient loads, and model performances were evaluated by standard assessment metrics—Root Mean Square Error (RMSE) and Correlation Coefficient (R). The residuals of predictive models were examined by the Durbin–Watson statistic. The results showed that MLP and LSTM persistently achieved better accuracy in predicting monthly TN and TP loads compared to GPR and ARIMA. In addition, GPR models achieved slightly better test RMSE score than ARIMA models while their correlation coefficients are much lower than ARIMA models. Moreover, MLP performed slightly better than LSTM in predicting monthly TP loads while LSTM slightly outperformed for TN loads. Furthermore, it was found that the optimizer and number of inputs didn’t show effects on the LSTM performance while they exhibited impacts on MLP outcomes. This study explores the capability of machine learning models to accurately predict nonlinearly fluctuating nutrient loads delivered to the Gulf of Mexico. Further efforts focus on improving the accuracy of forecasting using hybrid models which combine several machine learning models with superior predictive performance for nutrient fluxes throughout the MARB.

54 ENVIRONMENTAL SCIENCES

Exploring Spatiotemporal Relations Between Soil Moisture, Precipitation, and Streamflow for a Large Set of Watersheds Using Google Earth Engine

An understanding of streamflow variability and its response to changes in climate conditions is essential for water resource planning and management practices that will help to mitigate the impacts of extreme events such as floods and droughts on agriculture and other human activities. This study investigated the relationship between precipitation, soil moisture, and streamflow over a wide range of watersheds across the United States using Google Earth Engine (GEE). The correlation analyses disclosed a strong association between precipitation, soil moisture, and streamflow, however, soil moisture was found to have a higher correlation with the streamflow relative to precipitation. Results indicated different strength of the association depends on the watershed classes and lag times assessments. The perennial watersheds showed higher coherence compared to intermittent watersheds. Previous month precipitation and soil moisture have a stronger influence on the current month streamflow, particularly in the snow-dominated watersheds. Monthly streamflow forecasting models were developed using an autoregressive integrated moving average (ARIMA) and support vector machine (SVM). The results showed that the SVM model generally performed better than the ARIMA model. Overall streamflow forecasting model performance varied considerably among watershed classes, and perennial watersheds tend to exhibit better predictably compared to intermittent watersheds due to lower streamflow variability. The SVM models with precipitation and streamflow inputs performed better than those with streamflow input only. Results indicated that the inclusion of antecedent root-zone soil moisture improved the streamflow forecasting in most of the watersheds, and the largest improvements occurred in the intermittent watersheds. In conclusion, this work demonstrated that knowing the relationship between precipitation, soil moisture, and streamflow in different watershed classes will enhance the understanding of the hydrologic process and can be effectively utilized in improving streamflow forecasting for better satellite-based water resource management strategies.

Nazmus Sazib

Reliable statistics-based detection and investigation of anomalies in a SMART valve system

Reliable anomaly detection and diagnosis are critical for the safe operation of complex engineered systems. This study presents a unified framework that integrates statistical, model-based, and data-driven techniques for anomaly detection and investigation, demonstrated on SMART valve systems in hybrid energy applications. Four detection methods—mean deviation, seasonal extreme studentized deviate, ARIMA forecasting, and matrix profiling—were implemented and compared. Matrix profiling was particularly effective in revealing subtle deviations and hidden relationships among variables. Anomaly investigation was performed by analyzing variable-level and grouped signal profiles, with system topology incorporated to distinguish primary faults from propagated effects. Grouping signals by type enhanced interpretability, enabling accurate localization of anomalies across multi-dimensional datasets. Experimental results confirmed the framework's capability to consistently detect and isolate anomalies while providing actionable insights into system interdependencies. The proposed methodology offers a robust, interpretable, and scalable solution for condition monitoring, with potential applications in safety-critical domains such as nuclear energy, aerospace, and process industries.

ARIMA models

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING

A comprehensive framework to assess elemental mercury in the Department of Energy: A time series analysis

Objective: This study investigated whether seasonal categories affect airborne mercury concentrations in the U.S. Department of Energy operations. Methods: We conducted an initial assessment of the general variability of airborne elemental mercury time-weighted average (TWA) samples. Then, we performed a two-component time series analysis to determine whether long-term, cyclical temperature change patterns affect mercury concentrations. Results: Both ARIMA time series models demonstrated stationary, non-random means (χ² = 83.8, p < 0.001) and standard deviation (χ² = 55.8, p < 0.001) of mercury concentrations. Here, our results indicate that the seasonal factors did not influence mercury concentration. Conclusions: Our results demonstrate that mercury concentrations primarily emanate from operational activities, work practices, and/or transient environmental conditions rather than seasonal fluctuations.

Cannady, Ryan T. [Oak Ridge National Laboratory (O

Evaluating the Accuracy of Machine Learning Forecasts

To improve the accuracy of forecasting in machine learning, we must investigate multiple machine learning models and see how accurately they can predict values after training. We used seven machine learning models to try and get more accurate predictions. The models that were used were ARIMA, SES, MLP, CART, LightGBM, and XGBoost. We used a processed dataset from a Terminal at LAX that had the number of people traveling through terminal X every hour in March from 2015-2019. We trained our models with the dates March 6 - March 19 to predict the value for March 20th and the hours 6:00 am to 6:00 pm since those are the most popular traveling hours. By using the different models, we had varying results of accuracy when estimating the amount of people traveling through terminal X on March 20th. We know that machine learning models are helpful for forecasting and by seeing how accurately these models can predict, we can see how forecasting can be helpful for other issues. Using these methods, airports can use forecasting to predict the amount of people coming in and out and can use these predictions to prepare their resource management, operational efficiency, and overall passenger experience.

97 MATHEMATICS AND COMPUTING

Preliminary evidence for the influence of physiography and scale upon the autocorrelation function of remotely sensed data

Previously established results demonstrate that LANDSAT data are autocorrelated and can be described by a univariate linear stochastic process known as auto-regressive-integrated-moving-average model of degree 1, 0, 1 or ARIMA (1, 0, 1). This model has two coefficients of interest for interpretation phi(1) and theta(1). In a comparison of LANDSAT thematic mapper simulator (TMS) data and LANDSAT MSS data several results were established: (1) The form of the relatedness as described by this model is not dependent upon system look angle or pixel size. (2) The phi(1) coefficient increases with decreasing pixel size and increasing topographic complexity. (3) Changes in topography have a greater influence upon phi(1) than changes in land cover class. (4) The theta(1) seems to vary with the amount of atmospheric haze. These patterns of variation in phi(1) and theta(1) are potentially exploitable by the remote sensing community to yield stochastically independent sets of observations, characterize topography, and reduce the number of bytes needed to store remotely sensed data.

Labovitz, M. L.

Sources of variation in Landsat autocorrelation

Analysis of sixty-four scan lines representing diverse conditions across satellites, channels, scanners, locations and cloud cover confirms that Landsat data are autocorrelated and consistently follow an Arima (1,0,1) pattern. The AR parameter varies significantly with location and the MA coefficient with cloud cover. Maximum likelihood classification functions are considerably in error unless this autocorrelation is compensated for in sampling.

Craig, R. G.

Stochastic nature of Landsat MSS data

A multiple series generalization of the ARIMA models is used to model Landsat MSS scan lines as sequences of vectors, each vector having four elements (bands). The purpose of this work is to investigate if Landsat scan lines can be described by a general multiple series linear stochastic model and if the coefficients of such a model vary as a function of satellite system and target attributes. To accomplish this objective, an exploratory experimental design was set up incorporating six factors, four representing target attributes - location, cloud cover, row (within location), and column (within location) - and two factors representing system attributes - satellite number and detector bank. Each factor was included in the design at two levels and, with two replicates per treatment, 128 scan lines were analyzed. The results of the analysis suggests that a multiple AR(4) model is an adequate representation across all scan lines. Furthermore, the coefficients of the AR(4) model vary with location, particularly changes in physiography (slope regimes), and with percent cloud cover, but are insensitive to changes in system attributes.

Labovitz, M. L.

Stochastic Modeling of Airlines' Scheduled Services Revenue

Airlines' revenue generated from scheduled services account for the major share in the total revenue. As such, predicting airlines' total scheduled services revenue is of great importance both to the governments (in case of national airlines) and private airlines. This importance stems from the need to formulate future airline strategic management policies, determine government subsidy levels, and formulate governmental air transportation policies. The prediction of the airlines' total scheduled services revenue is dealt with in this paper. Four key components of airline's scheduled services are considered. These include revenues generated from passenger, cargo, mail, and excess baggage. By addressing the revenue generated from each schedule service separately, air transportation planners and designers are able to enhance their ability to formulate specific strategies for each component. Estimation results clearly indicate that the four stochastic processes (scheduled services components) are represented by different Box-Jenkins ARIMA models. The results demonstrate the appropriateness of the developed models and their ability to provide air transportation planners with future information vital to the planning and design processes.

Hamed, M. M.

Stochastic Modeling of Airlines' Scheduled Services Revenue

Airlines' revenue generated from scheduled services account for the major share in the total revenue. As such, predicting airlines' total scheduled services revenue is of great importance both to the governments (in case of national airlines) and private airlines. This importance stems from the need to formulate future airline strategic management policies, determine government subsidy levels, and formulate governmental air transportation policies. The prediction of the airlines' total scheduled services revenue is dealt with in this paper. Four key components of airline's scheduled services are considered. These include revenues generated from passenger, cargo, mail, and excess baggage. By addressing the revenue generated from each schedule service separately, air transportation planners and designers arc able to enhance their ability to formulate specific strategies for each component. Estimation results clearly indicate that the four stochastic processes (scheduled services components) are represented by different Box-Jenkins ARIMA models. The results demonstrate the appropriateness of the developed models and their ability to provide air transportation planners with future information vital to the planning and design processes.

Hamed, M. M.

Graph Representation Learning for Dengue Forecasting

In 2017, the largest recorded dengue outbreak in Sri Lanka’s history occurred. Since then, dengue has continued to threaten national health across Sri Lanka. The development of an effective Early Warning System (EWS) for dengue outbreaks is essential for Sri Lanka’s Ministry of Health to take preventative measures. We propose the use of Graph Neural Networks as EWS. Using earth observational data from NASAs global satellites and dengue incidence data from Sri Lanka s Ministry of Health, we developed a series of traditional and graph representation EWS to forecast Dengue cases across Sri Lanka’s 25 districts between 2013 and 2022. We demonstrate empirically that Graph Neural Networks which incorporate spatiotemporal relations significantly outperform traditional EWS such as Autoregressive Integrated Moving Average (ARIMA), Random Forest, and Long Short-Term Memory (LSTM). Our source code is available on GitHub and will be provided in the final submission.

Graph Neural Networks

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES

Measurement of electrical characteristics of Ge-based diodes with thermal GeO2 films exposed to controlled humidity conditions

We investigated the impact of the interaction between gas-phase water molecules and thermally oxidized GeO2 on Ge on the performance of metal–oxide–semiconductor (MOS) structures. A vacuum-integrated setup was developed to form MOS diodes with a GeO2/Ge structure that had been exposed to controlled humidity conditions and to subsequently measure their electrical characteristics in situ. The capacitance–voltage (C–V) curves exhibited a significant negative shift, indicating the generation of positive charges at the GeO2/Ge interface when the GeO2 surface was exposed to humidity levels above approximately 1%. According to a previous study using electron spectroscopy, this threshold corresponds to the humidity level at which molecular water begins to grow on a GeO2/Ge structure. It is likely that gas-phase water molecules infiltrating the GeO2 film bind to local OH sites via hydrogen bonding, leading to the formation of positive fixed charges at the GeO2/Ge interface.

Sano, Shuto

Phonon excitations in Eu 2 Ir 2 O 7 probed by inelastic x-ray scattering

The study of phonon dynamics and its interplay with magnetic ordering is crucial for understanding the unique quantum phases in the pyrochlore iridates. Here, through inelastic x-ray scattering on a single crystal sample of the pyrochlore iridate Eu 2 Ir 2 O 7 , we map out the phonon excitation spectra in Eu 2 Ir 2 O 7 and compare them with the theoretical phonon spectra calculated using the density functional theory. Further, possible phonon renormalization across the magnetic long-range order transition is observed in our experiments, which is consistent with the results of the previous Raman scattering experiments.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Collinear ferromagnetism with reduced moment length in kagome magnet Nd 3 ⁢Ru 4 ⁢Al 12

Here, we determine the magnetic ground state of the kagome lattice magnet Nd 3 ⁢Ru 4⁢ Al 12 by single-crystal neutron diffraction, supported by experiments with polarized neutrons. We identify this material as a collinear ferromagnet (“hex-FM”) with uniform moment length and ordering vector 𝑸 = 0, in contrast to a previous, seminal report that proposed unequal moment lengths on two Nd sites, here called the “ortho-FM” state. Our analysis of the flipping ratio in polarized neutron scattering is consistent with the hex-FM state. The results provide a microscopic basis for understanding the large fluctuation-induced Hall and Nernst responses near 𝑇 C ≈ 41K, as previously reported for Nd 3 ⁢Ru 4 ⁢Al 12 .

RKKY interaction