Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “support vector regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Automatic Waveform Quality Control for Surface Waves Using Machine Learning

Surface-wave seismograms are widely used by researchers to study Earth’s interior and earthquakes. To extract information reliably and robustly from a suite of surface waveforms, the signals require quality control screening to reduce artifacts from signal complexity and noise. This process has usually been completed by human experts labeling each waveform visually, which is time consuming and tedious for large data sets. We explore automated approaches to improve the efficiency of waveform quality control processing by investigating logistic regression, support vector machines, K-nearest neighbors, random forests (RF), and artificial neural networks (ANN) algorithms. To speed up signal quality assessment, we trained these five machine learning (ML) methods using nearly 400,000 human-labeled waveforms. The ANN and RF models outperformed other algorithms and achieved a test accuracy of 92%. We evaluated these two best-performing models using seismic events from geographic regions not used for training. The results show that the two trained models agree with labels from human analysts but required only 0.4% of the time. Although the original (human) quality assignments assessed general waveform signal-to-noise, the ANN or RF labels can help facilitate detailed waveform analysis. Our investigations demonstrate the capability of the automated processing using these two ML models to reduce outliers in surface-wave-related measurements without human quality control screening.

58 GEOSCIENCES↗

Assessment of Machine Learning for Ultrasonic Nondestructive Evaluation of Alkali–Silica Reaction in Concrete

Alkali–silica reaction (ASR) is a type of material degradation in concrete structures that leads to concrete cracking and rebar corrosion, thereby reducing the material’s structural integrity and the overall structure’s lifetime and raising safety concerns. Ultrasonic nondestructive evaluation (NDE) has been proven to be a valuable technique for assessing concrete properties and monitoring ASR progression in concrete. However, the deployment and analysis of ultrasonic NDE and its data requires specialized expertise, often relying on the engineer’s subjective interpretation. With the surge in computational power, artificial intelligence (AI) and machine learning (ML) algorithms have become popular in automating NDE data analysis. Various industrial sectors are increasingly adopting ML algorithms for NDE data analysis with a growing emphasis on AI–assisted automation. Regulatory agencies are also preparing for this technological shift, anticipating corresponding revisions in standards. Thus, there is an urgent need to identify the capabilities and limitations of current ML technologies for the evaluation of concrete material properties and damage status. Furthermore, the effects of various factors on ML model performance must be thoroughly investigated. The study summarized herein evaluated the effectiveness of two ML models (i.e., support vector regression (SVR) and deep neural network (DNN)) in predicting concrete material damage induced by ASR based on the long-term ultrasonic monitoring data. Four distinct concrete specimens were cast with artificially induced ASR, and over a period exceeding 500 days, ultrasonic signals and expansion data were continuously collected. For the SVR model, wave velocity and 12 other wave features were extracted from the ultrasonic signals, with 6 out of 13 features selected as input for the model. Different combinations of training and testing datasets were designed to explore factors influencing prediction performance, including the range of data within training and testing sets, in addition to various signal preprocessing methodologies. These findings suggest the importance of using a training dataset with a broader data range compared with testing datasets for improved model performance alongside consistent signal preprocessing across datasets.

36 MATERIALS SCIENCE↗

Integration of LIBS with Machine Learning for Real-Time Monitoring of Feedstock in H 2 Gasification Applications

This project, funded by the U.S. Department of Energy (DOE) – Office of Fossil Energy under Award Number DE-FE0032177, aimed to assess the feasibility of an integrated Laser-Induced Breakdown Spectroscopy (LIBS) system with advanced machine learning (ML) models for real-time characterization and potential control of hydrogen gasifiers running on waste materials as feedstocks. This was a multidisciplinary effort that encompassed the acquisition and standardized analysis of individual and blended feedstocks—comprising biomass, coal waste, and plastic waste, followed by the development of a dynamic LIBS bench system for material sample analysis and development of predictive ML models. Comprehensive laboratory testing enabled the creation of a robust elemental dataset that served as the foundation for ML model training. Techniques such as Random Forest, Gradient Boosting, Support Vector Regression, and Neural Networks were employed to predict key feedstock properties, including higher heating value (HHV), moisture content, thermal conductivity, and ash composition with high accuracy. The results were validated against experimental data and demonstrated strong potential for real-time application in gasifier control systems. The project concluded with a study on the integration of the LIBS+ML approach for gasifier control and a techno-economic analysis of the implementation of the approach into hydrogen (H 2 ) gasification systems. Dissemination of results was carried out at a DOE meeting. This work establishes a scalable framework for automated, in-line feedstock quality assessment, offering significant implications for process optimization and emissions reduction in hydrogen production.

01 COAL, LIGNITE, AND PEAT↗

Projecting the Thermal Response in a HTGR-Type System during Conduction Cooldown Using Graph-Laplacian Based Machine Learning

Accurate prediction of an off-normal event in a nuclear reactor is dependent upon the availability of sensory data, reactor core physical condition, and understanding of the underlying phenomenon. This work presents a method to project the data from some discrete sensory locations to the overall reactor domain during conduction cooldown scenarios similar to High Temperature Gas-cooled Reactors (HTGRs). The existing models for conductive cooldown in a heterogeneous multi-body system, such as an assembly of prismatic blocks or pebble beds relies on knowledge of the thermal contact conductance, requiring significant knowledge of local thermal contacts and heat transport possibilities across those contacts. With a priori knowledge of bulk geometry features and some discrete sensors, a machine learning approach was devised. The presented work uses an experimental facility to mimic conduction cooldown with an assembly of 68 cylindrical rods initially heated to 1200 K. High-fidelity temperature data were collected using an infrared (IR) camera to provide training data to the model and validate the predicted temperature data. The machine learning approach used here first converts the macroscopic bulk geometry information into Graph-Laplacian, and then uses the eigenvectors of the Graph-Laplacian to develop Kernel functions. Support vector regression (SVR) was implemented on the obtained Kernels and used to predict the thermal response in a packed rod assembly during a conduction cooldown experiment. The usage of SVR modeling differs from most models today because of its representation of thermal coupling between rods in the core. When trained with thermographic data, the average normalized error is less than 2% over 400 s, during which temperatures of the assembly have dropped by more than 500 K. The rod temperature prediction performance was significantly better for rods in the interior of the assembly compared to those near the exterior, likely due to the model simplification of the surroundings.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning

Stream temperature (Ts) is an important water quality parameter that affects ecosystem health and human water use for beneficial purposes. Accurate Ts predictions at different spatial and temporal scales can inform water management decisions that account for the effects of changing climate and extreme events. In particular, widespread predictions of Ts in unmonitored stream reaches can enable decision makers to be responsive to changes caused by unforeseen disturbances. In this study, we demonstrate the use of classical machine learning (ML) models, support vector regression and gradient boosted trees (XGBoost), for monthly Ts predictions in 78 pristine and human-impacted catchments of the Mid-Atlantic and Pacific Northwest hydrologic regions spanning different geologies, climate, and land use. The ML models were trained using long-term monitoring data from 1980–2020 for three scenarios: (1) temporal predictions at a single site, (2) temporal predictions for multiple sites within a region, and (3) spatiotemporal predictions in unmonitored basins (PUB). In the first two scenarios, the ML models predicted Ts with median root mean squared errors (RMSE) of 0.69–0.84 °C and 0.92–1.02 °C across different model types for the temporal predictions at single and multiple sites respectively. For the PUB scenario, we used a bootstrap aggregation approach using models trained with different subsets of data, for which an ensemble XGBoost implementation outperformed all other modeling configurations (median RMSE 0.62 °C).The ML models improved median monthly Ts estimates compared to baseline statistical multi-linear regression models by 15–48% depending on the site and scenario. Air temperature was found to be the primary driver of monthly Ts for all sites, with secondary influence of month of the year (seasonality) and solar radiation, while discharge was a significant predictor at only 10 sites. The predictive performance of the ML models was robust to configuration changes in model setup and inputs, but was influenced by the distance to the nearest dam with RMSE <1 °C at sites situated greater than 16 and 44 km from a dam for the temporal single site and regional scenarios, and over 1.4 km from a dam for the PUB scenario. Our results show that classical ML models with solely meteorological inputs can be used for spatial and temporal predictions of monthly Ts in pristine and managed basins with reasonable (<1 °C) accuracy for most locations.

54 ENVIRONMENTAL SCIENCES↗

Analyzing Double Delays at Newark Liberty International Airport

When weather or congestion impacts the National Airspace System, multiple different Traffic Management Initiatives can be implemented, sometimes with unintended consequences. One particular inefficiency that is commonly identified is in the interaction between Ground Delay Programs (GDPs) and time based metering of internal departures, or TMA scheduling. Internal departures under TMA scheduling can take large GDP delays, followed by large TMA scheduling delays, because they cannot be easily fitted into the overhead stream. In this paper we examine the causes of these double delays through an analysis of arrival operations at Newark Liberty International Airport (EWR) from June to August 2010. Depending on how the double delay is defined between 0.3 percent and 0.8 percent of arrivals at EWR experienced double delays in this period. However, this represents between 21 percent and 62 percent of all internal departures in GDP and TMA scheduling. A deep dive into the data reveals that two causes of high internal departure scheduling delays are upstream flights making up time between their estimated departure clearance times (EDCTs) and entry into time based metering, which undermines the sequencing and spacing underlying the flight EDCTs, and high demand on TMA, when TMA airborne metering delays are high. Data mining methods (currently) including logistic regression, support vector machines and K-nearest neighbors are used to predict the occurrence of double delays and high internal departure scheduling delays with accuracies up to 0.68. So far, key indicators of double delay and high internal departure scheduling delay are TMA virtual runway queue size, and the degree to which estimated runway demand based on TMA estimated times of arrival has changed relative to the estimated runway demand based on EDCTs. However, more analysis is needed to confirm this.

traffic management advisor↗

Space-Borne Cloud-Native Satellite-Derived Bathymetry (SDB) Models Using ICESat-2 And Sentinel-2

Shallow nearshore coastal waters provide a wealth of societal, economic and ecosystem services, yet their topographic structure is poorly mapped due to a reliance upon expensive and time intensive methods. Space‐borne bathymetric mapping has helped address these issues, but has remained largely dependent upon in situ measurements. Here we fuse ICESat‐2 lidar data with Sentinel‐2 optical imagery, within the Google Earth Engine cloud platform, to create openly available spatially continuous high‐resolution bathymetric maps at regional‐to‐national scales in Florida, Crete and Bermuda. ICESat‐2 bathymetric classified photons are used to train three Satellite Derived Bathymetry (SDB) methods, including Lyzenga, Stumpf and Support Vector Regression algorithms. For each study site the Lyzenga algorithm yielded the lowest RMSE (approx. 10‐15%) when compared with validation data. We demonstrate a means of using ICESat‐2 for both model calibration and validation, thus cementing a pathway for fully space‐borne estimates of nearshore bathymetry in shallow, clear water environments.

N. Thomas↗

Comparative Analysis of Empirical and Machine Learning Models for Chla Extraction Using Sentinel-2 and Landsat OLI Data: Opportunities, Limitations, and Challenges

Remote retrieval of near-surface chlorophyll-a (Chla) concentration in small inland waters is challenging due to substantial optical interferences of various water constituents and uncertainties in the atmospheric correction (AC) process. Although various algorithms have been developed to estimate Chla from moderate-resolution terrestrial missions (∼10–60 m), the production of both accurate distribution maps and time series of Chla has proven challenging, limiting the use of remote analyses for lake monitoring. Here, we develop a support vector regression (SVR) model, which uses satellite-derived remote-sensing reflectance spectra () from Sentinel-2 and Landsat-8 images as input for Chla retrieval in a representative eutrophic prairie lake, Buffalo Pound Lake (BPL), Saskatchewan, Canada. Validated against in situ Chla from seven ice-free seasons (N ∼ 200; 2014–2020), the SVR model outperformed both locally tuned, -fed empirical models (Normalized Difference Chlorophyll Index, 2- and 3-band, and OC3) and Mixture Density Networks (MDNs) by 15–65%, while exhibiting comparable performance to a locally trained MDN, with an error of ∼35%. Comparison of Chla retrieval models, AC processors (iCOR, ACOLITE), and radiometric products (Rayleigh-corrected, surface, and top-of-atmosphere reflectance) showed that the best Chla maps and optimal time series (up to 100 mg m−3) were produced using a coupled SVR-iCOR system.

algal blooms↗

Physics-Infused AI/ML Based Digital-Twin Framework for Flow-Induced-Vibration Damage Prediction in a Nuclear Reactor Heat Exchanger

This report summarizes some of the ongoing work related to the development of an expert-elicitation-digital-twin framework for real time damage state prediction in heat exchanger components of a nuclear reactor. The framework is targeted towards predicting damage associated with coupled low cycle fatigue (associated with regular heat-up, cool-down and power operation transients) and high cycle fatigue (associated with flow induced vibration transients). The overall framework will be based on a NoSQL based database, physics-infused-geometry-dependent virtual-sensor data, different AI/ML techniques-based data-driven-predictive-model applications (Apps) and real-time plant sensor measurements available through few existing sensors. Towards this overall goal, this report updates some of the ongoing work, such as on implementation of a NoSQL Database (such as MongoDB), FE based heat transfer analysis of a heat exchanger (e.g. of a PWR steam generator) for generating geometry-dependent virtual sensor data and evaluation of various AI/ML models such as based on multivariate linear regression, ensembled decision-tree based Random-Forest and Gradient-Boosting regression and high-dimensional-kernel-function-transformation based Support-Vector-Machine regression models. The AI/ML models were evaluated for predicting multi-time-series thermal states at thousands of 3D point-clouds

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Advancing river corridor science beyond disciplinary boundaries with an inductive approach to catalyse hypothesis generation

Abstract A unified conceptual framework for river corridors requires synthesis of diverse site‐, method‐ and discipline‐specific findings. The river research community has developed a substantial body of observations and process‐specific interpretations, but we are still lacking a comprehensive model to distill this knowledge into fundamental transferable concepts. We confront the challenge of how a discipline classically organized around the deductive model of systematically collecting of site‐, scale‐, and mechanism‐specific observations begins the process of synthesis. Machine learning is particularly well‐suited to inductive generation of hypotheses. In this study, we prototype an inductive approach to holistic synthesis of river corridor observations, using support vector machine regression to identify potential couplings or feedbacks that would not necessarily arise from classical approaches. This approach generated 672 relationships linking a suite of 157 variables each measured at 62 locations in a fifth order river network. Eighty four percent of these relationships have not been previously investigated, and representing potential (hypothetical) process connections. We document relationships consistent with current understanding including hydrologic exchange processes, microbial ecology, and the River Continuum Concept, supporting that the approach can identify meaningful relationships in the data. Moreover, we highlight examples of two novel research questions that stem from interpretation of inductively‐generated relationships. This study demonstrates the implementation of machine learning to sieve complex data sets and identify a small set of candidate relationships that warrant further study, including data types not commonly measured together. This structured approach complements traditional modes of inquiry, which are often limited by disciplinary perspectives and favour the careful pursuit of parsimony. Finally, we emphasize that this approach should be viewed as a complement to, rather than in place of, more traditional, deductive approaches to scientific discovery.

54 ENVIRONMENTAL SCIENCES↗

Estimating Dust and Water Ice Content of the Martian Atmosphere From THEMIS Data

Researchers at JPL and Arizona State University conducted a comparative study of three candidate algorithms for estimating components of the Martian atmosphere, using raw (uncalibrated) data collected by the Thermal Emission Imaging System (THEMIS). THEMIS is an instrument onboard the Mars Odyssey spacecraft that acquires image data in five visible and nine infrared (IR) wavelength bands. The algorithms under study used data collected from eight of the nine IR bands to estimate the dust and water ice content of the atmosphere. Such an algorithm could be used in onboard data processing to trigger other algorithms that search for features of scientific interest and to reduce the volume of data transmitted to Earth. The algorithms studied were based on regression models. In the study, the optical depths estimated by these algorithms were compared with optical depths estimated in ground-based processing using fully calibrated data from both THEMIS and the Thermal Emission Spectrometer (TES). TES is an instrument onboard the Mars Global Surveyor spacecraft that also observes the planet at infrared wavelengths, but at a lower spatial resolution than THEMIS does. Of the algorithms studied, the one that performed best was based on a Gaussian Support Vector Machine regression model. The test results indicated that this algorithm, operating on the raw data, had error rates that were within the uncertainty associated with the estimates obtained by the groundbased analysis of the fully calibrated data. This level of fidelity demonstrates that these algorithms are sufficiently accurate for use in an onboard setting.

Bandfield, Joshua↗

A machine learning approach to water quality forecasts and sensor network expansion: Case study in the Wabash River Basin, United States

Abstract Midwestern cities require forecasts of surface nitrate loads to bring additional treatment processes online or activate alternative water supplies. Concurrently, networks of nitrate monitoring stations are being deployed in river basins, co‐locating water quality observations with established stream gauges. However, tools to evaluate the future value of expanded networks to improve water quality forecasts remains challenging. Here, we construct a synthetic data set of stream discharge and nitrate for the Wabash River Basin—one of the United States’ most nutrient polluted basins—using the established Agro‐IBIS and THMB models. Synthetic data enables rapid, unbiased and low‐cost assessment of potential sensor placements to support management objectives, such as near‐term forecasting. Using the synthetic data, we established baseline 1‐day forecasts for surface water nitrate at 12 cities in the basin using support vector machine regression (SVMR; RMSE 0.48–3.3 ppm). Next, we used the SVMRs to evaluate the improvement in forecast performance associated with deployment of additional nitrate sensors. We identified the optimal sensor placement to improve forecasts at each city, and the relative value of sensors at each candidate location. Finally, we assessed the co‐benefit realized by other cities when a sensor is deployed to optimize a forecast at one city, finding significant positive externalities in all cases. Ultimately, our study explores the potential for machine learning to make near‐term predictions and critically evaluate the improvement realized by expanding a monitoring network. While we use nitrate pollution in the Wabash River Basin as a case study, this approach could be readily applied to any problem where the future value of sensors and network design are being evaluated.

54 ENVIRONMENTAL SCIENCES↗

Risk assessment of engineering diseases of embankment–bridge transition section for railway in permafrost regions

Abstract The embankment–bridge transition section (EBTS) is one of the zones where railway diseases occur frequently in permafrost regions. Disease risk assessment of EBTSs can provide guidance for maintenance. In this study, considering the engineering geological conditions, climate characteristics, and embankment structure types along the Qinghai–Tibet Railway (QTR) as well as based on the disease inventory of the QTR from 2010 to 2019, the logistic regression (LR), support vector machine (SVM), and combination‐weight‐based gay relation analysis (GRA) were used for disease risk assessment of the EBTSs along the QTR in permafrost regions. The results indicate that the LR and SVM models have a better capability for EBTS disease prediction than the GRA model, and the SVM model can select more disease samples in relatively larger regions than the LR model. Based on the SVM and LR models, the risk level of EBTSs is divided into four classes: low‐ (29.9%), moderate‐ (39.6%), high‐ (22.1%), and very high (8.4%) risk. Finally, we selected 272 EBTSs in high‐ and very‐high‐risk classes for key observation during the maintenance of the QTR in permafrost regions. This study provides a reference for the risk assessment of railways built in permafrost regions using data‐driven methods.

Zhang, Saize↗

Situational awareness-enhancing community-level load mapping with opportunistic machine learning

Motivated by present and forthcoming challenges in the adoption and integration of distributed renewable energy, we develop a machine learning (ML) approach that builds short-fuse mappings connecting the occasionally-unobservable true load in one target community with information-rich signals collected from relatively more instrumented reference communities. Our setting is inspired by and tailored to target communities with significant unobservable behind-the-meter solar generation, where true load (a relatively well-behaved quantity of interest to grid operators) is hard to discern during daytime due to insufficient instrumentation and/or privacy reasons, but that can be related to reference communities with low unobservable distributed variable generation or with sufficient instrumentation. The developed mapping, herein realized with Support Vector Machine regression, is built using nighttime data from all communities, when their distributed generation is low or zero. Our ML algorithm opportunistically learns to correlate signals of interest and then is operationally used the next day to shed light into target community load evolution. The mapping is subsequently rebuilt, rolling its short-fuse scope perpetually forward in time. Here, we demonstrate the efficacy of our approach on nine synthetically generated topologies and associated timeseries stemming from real-world data, on which we observe cumulative error performance that yields lower than 10% and 15% daily-averaged mean absolute percentage errors in target community load estimation on more than about 75% and 90% of days, respectively, in multiple yearly evaluations that shed light on long-term performance also under seasonal and one-off effects. The proposed ML-powered methodology can offer grid operators much-improved visibility into a previously obscure space and can also serve as an additional source of information in broader, multi-modal solar disaggregation solutions.

14 SOLAR ENERGY↗

A machine learning approach to determine the elastic properties of printed fiber-reinforced polymers

This work focuses on the simultaneous determination of the elastic constants and the fiber orientation state for a short fiber-reinforced polymer composite by performing a minimum of experimental tests. Here we introduce a methodology that enables the inverse determination of fiber orientation state and the in-situ polymer properties by performing tensile tests at the composite coupon level. We demonstrate the approach for the extrusion deposition additive manufacturing (EDAM) process to illustrate one application of the methodology, but the development is such that it can be applied to short fiber-reinforced polymer (SFRP) systems processed via other methods. Currently, developing composites additive manufacturing digital twins require extensive material characterization. In particular, the mechanical characterization of the orthotropic elastic properties of a composite involves extensive sample preparation and testing, therefore the elasticity tensor is generally populated using a micromechanics model. This, however, requires measuring the fiber orientation state in addition to knowing the constituent material properties. Experimentally measuring the fiber orientation state can be tedious and time consuming. Further, optical methods are limited to resolving the orientation of cylindrical fibers or cluster of non-cylindrical fibers, and computed tomography (CT) methods scan regions of volume that are much smaller than a full printed bead. Therefore, we propose a methodology, accelerated by machine learning, to identify the anisotropic mechanical properties and fiber orientation state at the same time. Early results show that inference of the fiber orientation and composite properties is possible with as few as three tensile tests. Our results show that a combination of the choice of the micromechanics model and reliable set of experiments can yield the nine elastic constants, as well as, the fiber orientation state.

36 MATERIALS SCIENCE↗

Applying NIR and MIR spectroscopy for C and soil property prediction in northern cold-region ecosystems. Which approach works better?

Here, developing reliable predictions of soil attributes is necessary to understand northern cold-region climate-soil feedback. Calibration models using near-infrared (NIR) and mid-infrared (MIR) spectroscopy were developed to predict eight commonly measured soil properties for 119 soil samples representing a range of vegetation types, parent materials, and soil types spanning >23° of latitude from southeast Alaska to the Canadian high Arctic. In order to obtain a more accurate prediction, this study compared the performance of linear and non-linear calibration techniques, including lasso regression (Lasso), support vector machine (SVM), random forest (RF) and classic partial least squares (PLS) to predict different soil properties of these soils. Comparing the four models, we noticed that their performance was quite similar for MIR overall, while NIR achieved better results with a PLS model for our dataset. PLS coupled with MIR showed a better performance for soil parameters, such as total organic carbon (TOC), total nitrogen (TN), cation exchange capacity (CEC) and clay (R-squared of 0.9, 0.81, 0.80, and 0.84) when compared with NIR (R-squared of 0.85, 0.72, 0.81 and 0.68). However, using either MIR or NIR spectroscopy, PLS predictions for bulk density (BD) and sand content were not accurate. The variable importance analysis based on the PLS model successfully estimated the relative contribution of wavelengths influencing soil property predictions most. Overall, TOC, TN, CEC and clay mineral predictions are closely related to the occurrence of specific spectral bands in the MIR region. For example, wavelengths at 2978 and 1761 cm -1 for TOC and TN, as well as at 3064 cm -1 for CEC, were selected as the most influential predictor variables. We demonstrated that MIR spectroscopy is a powerful tool for more extensive monitoring in soils of the northern cold climate region; however, NIR could be utilized for rapid estimates when the highest accuracy is not essential.

54 ENVIRONMENTAL SCIENCES↗