Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Altimeter waveform software design

Techniques are described for preprocessing raw return waveform data from the GEOS-3 radar altimeter. Topics discussed include: (1) general altimeter data preprocessing to be done at the GEOS-3 Data Processing Center to correct altimeter waveform data for temperature calibrations, to convert between engineering and final data units and to convert telemetered parameter quantities to more appropriate final data distribution values: (2) time "tagging" of altimeter return waveform data quantities to compensate for various delays, misalignments and calculational intervals; (3) data processing procedures for use in estimating spacecraft attitude from altimeter waveform sampling gates; and (4) feasibility of use of a ground-based reflector or transponder to obtain in-flight calibration information on GEOS-3 altimeter performance.

Hayne, G. S.↗

GNET2: an R package for constructing gene regulatory networks from transcriptomic data

Abstract Motivation The Gene Network Estimation Tool (GNET) is designed to build gene regulatory networks (GRNs) from transcriptomic gene expression data with a probabilistic graphical model. The data preprocessing, model construction and visualization modules of the original GNET software were developed on different programming platforms, which were inconvenient for users to deploy and use. Results Here, we present GNET2, an improved implementation of GNET as an integrated R package. GNET2 provides more flexibility for parameter initialization and regulatory module construction based on the core iterative modeling process of the original algorithm. The data exchange interface of GNET2 is handled within an R session automatically. Given the growing demand for regulatory network reconstruction from transcriptomic data, GNET2 offers a convenient option for GRN inference on large datasets. Availability and implementation The source code of GNET2 is available at https://github.com/jianlin-cheng/GNET2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Investigation of LANDSAT D Thematic Mapper geometric performance: Line to line and band to band registration

The geometric accuray of LANDSAT TM raw data of Toulouse (France) raw data of Mississippi, and preprocessed data of Mississippi was examined using a CDC computer. Analog images were restituted on the VIZIR SEP device. The methods used for line to line and band to band registration are based on automatic correlation techniques and are widely used in automated image to image registration at CNES. Causes of intraband and interband misregistration are identified and statistics are given for both line to line and band to band misregistration.

Begni, G.↗

A survey of automated remote sensing for agriculture

The state-of-the-art of the technology available to make remote sensing crop production estimates is reviewed with reference to several past and present research projects. In particular, attention is given to Landsat data acquisition, registration and preprocessing, data transformation, data modeling, proportion estimation, and labeling. Development stage models and crop condition models are briefly characterized, and areas where further research is needed are identified.

Hall, F. G.↗

User microprogrammable processors for high data rate telemetry preprocessing

The use of microprogrammable processors for the preprocessing of high data rate satellite telemetry is investigated. The following topics are discussed along with supporting studies: (1) evaluation of commercial microprogrammable minicomputers for telemetry preprocessing tasks; (2) microinstruction sets for telemetry preprocessing; and (3) the use of multiple minicomputers to achieve high data processing. The simulation of small microprogrammed processors is discussed along with examples of microprogrammed processors.

Pugsley, J. H.↗

A generalized machine learning workflow to visualize mechanical discontinuity

Accurate detection and mapping of mechanical discontinuity in materials has widespread industrial and research applications. Herein, we developed a generalized machine-learning framework for visualizing single mechanical discontinuity embedded in material of any composition, velocity, density, porosity, and size with limited data. The proposed visualization of discontinuity requires accurate estimations of the length, location, and orientation of the embedded discontinuity by processing multipoint wave-transmission measurements. k-Wave simulator is used to create a large dataset of elastic waveforms recorded during multi-point wave-transmission measurements through materials containing single mechanical discontinuity. k-Wave simulator considers the wave attenuation, dispersion, and mode conversion in wave motion. Discrete wavelet transform (DWT) and statistical feature extraction are essential for data preprocessing prior to the data-driven model development. DWT also minimizes the effect of noise. Using hyper-parameter tuning and cross validation, gradient boosting regression can visualize the mechanical discontinuity with an accuracy of 0.85, in terms of coefficient of determination. A double-layered neural network-based regression has better performance with an accuracy of 0.95. Use of convolutional neural network converts the predictive task from a waveform processing to an image processing problem. Convolutional neural network achieved a generalization performance of 0.91. The proposed generalized workflow requires robust simulation of wave propagation, signal processing, feature engineering, and model evaluation. Sensors closest to the source and those located opposite the source are the most significant for the desired visualization. Notably, the sensors closest to the source capture the non-linear associations, whereas the sensor on the border opposite to the source capture the linear associations between the measured waveforms and the properties of the mechanical discontinuity.

42 ENGINEERING↗

Framework For Spatial Agricultural Crop Yield Prediction Model Development

This framework was developed to provide data preprocessing for spatiotemporal agricultural yield data and remote sensing data for modelling using artificial neural networks (ANNs) to predict subfield crop yield estimates. The software includes methods to train, validate, and test ANN models. It also include methods to infer on new remote sensing data.

Griffel, LloydM.↗

Survey of Use Cases and Scenarios on the Open Energy Data Initiative Solar Systems Integration (OEDI SI) Platform

The Open Energy Data Initiative Solar Systems Integration (OEDI SI) Data and Modeling Platform offers a comprehensive set of use cases tailored for power systems analysis. Each use case is centered around a specific power system analysis problem, supported by composite input data and reference algorithms. These composite input datasets are meticulously assembled using OEDI SI's data preprocessing tools, which integrate raw data from various sources. The primary objectives of the OEDI SI Platform include facilitating access to composite input data through widely accepted input/output formats and verified results. This accessibility enables power system network researchers and developers to validate their algorithms and showcase their applications' capabilities to the broader community. Moreover, the platform strives to promote reproducible, robust, replicable, and generalizable solar systems integration research.

14 SOLAR ENERGY↗

Data processing 2: Advancements in large scale data processing systems for remote sensing

The development of large scale data processing systems for remote sensing is studied by evaluating: (1) the suitability of several sensor types with regard to producing data required for multispectral machine analysis; (2) various types of data preprocessing necessary to prepare such data for analysis; and (3) transfer of machine processing techniques for earth resources data to user community.

Landgrebe, D. A.↗

Data processing assessment for the Lunar Geoscience Observer imaging spectrometer

On the Lunar Geoscience Observer project, a Visible and Infrared Mapping Spectrometer instrument has been proposed. This instrument will have science data input rates in the hundreds of kilobits per second (kbps) and an average telemetry output data rate of 4 kbps. Techniques that can be used to reduce the throughput of the instrument are editing, summing and averaging, data compression, data preprocessing, pattern recognition and snapshot data taking. Due to instrument limitations in the buffer memory size and processing speeds, a careful selection of the available techniques must be made.

Irigoyen, R. E.↗

Arithmetic Primitives for Efficient Neuromorphic Computing

Neuromorphic computing is steadily gaining popularity in many scientific and engineering disciplines. However, one of the biggest problems that has prevented widespread usage of neuromorphic computing is the lack of efficient encoding methods. Traditional encoding methods such as binning, rate encoding, and temporal encoding are based on unary encoding and generate a large number of spikes for certain applications, making them less energy efficient. Lack of better encoding methods has also prevented preprocessing operations from being carried out on neuromorphic computers. As a result, over 99% of the time can be spent on data preprocessing and data transfer operations in some cases, leading to an inefficient workflow. In this paper, we present preliminary results that would enable us to efficiently encode data and perform basic arithmetic operations on neuromorphic computers. First, we present a neuromorphic approach for the two’s complement encoding of numbers and leverage it to devise addition and multiplication circuits, which could be used in preprocessing operations on neuromorphic computers. We test our approach on the SuperNeuroMAT simulator. Our results indicate that two’s complement is a highly efficient encoding method in terms of time, space, and energy complexity and that the addition and multiplication circuits produce accurate results on two numbers having arbitrary precision.

Wurm, Ahna↗

The preprocessing of multispectral data. II

It is pointed out that a correction of atmospheric effects is an important requirement for a full utilization of the possibilities provided by preprocessing techniques. The most significant characteristics of original and preprocessed data are considered, taking into account the solution of classification problems by means of the preprocessing procedure. Improvements obtainable with different preprocessing techniques are illustrated with the aid of examples involving Landsat data regarding an area in Colorado.

Quiel, F.↗

Recent developments with the ORSER system

Additions to the ORSER remote sensing data processing package are described. The ORSER package consists of about 35 individual programs that are grouped into preprocessing, data analysis, and display subsystems. Additional data formats and data management, data transformation, and geometric correlation programs were supplemented to the preprocessing subsystem. Enhancements to the data analysis techniques include a maximum likelihood classifier (MAXCLASS) and a new version of the STATS program which makes delineation of training areas easier and allows for detection of outlier points. Ongoing developments are also described.

Baumer, G. M.↗

A data-centric weak supervised learning for highway traffic incident detection

Using the data from loop detector sensors for near-real-time detection of traffic incidents on highways is crucial to averting major traffic congestion. While recent supervised machine learning methods offer solutions to incident detection by leveraging human-labeled incident data, the false alarm rate is often too high to be used in practice. Specifically, the inconsistency in the human labeling of the incidents significantly affects the performance of supervised learning models. To that end, we focus on a data-centric approach to improve the accuracy and reduce the false alarm rate of traffic incident detection on highways. We develop a weak supervised learning workflow to generate high-quality training labels for the incident data without the ground truth labels, and we use those generated labels in the supervised learning setup for final detection. This approach comprises three stages. First, we introduce a data preprocessing and curation pipeline that processes traffic sensor data to generate high-quality training data through leveraging labeling functions, which can be domain knowledge-related or simple heuristic rules. Second, we evaluate the training data generated by weak supervision using three supervised learning models-random forest, k-nearest neighbors, and a support vector machine ensemble-and long short-term memory classifiers. The results show that the accuracy of all of the models improves significantly after using the training data generated by weak supervision. Third, we develop an online real-time incident detection approach that leverages the model ensemble and the uncertainty quantification while detecting incidents. Finally, we show that our proposed weak supervised learning workflow achieves a high incident detection rate (0.90) and low false alarm rate (0.08).

97 MATHEMATICS AND COMPUTING↗

Experiments in data collection technology using satellites

A variety of techniques potentially useful to data collection have been tested. An automatic data collection platform with a minicomputer collects and preprocesses data, then sends desired information when interrogated through a communication satellite. Position surveillance by tone-code ranging through communication satellites is automatic, real time and accurate. Emergency medical data transmissions from ambulances to hospitals can be extended to rural and remote areas by direct satellite links. A small platform can send emergency-related data through a satellite while the satellite is routinely relaying powerful communication signals. A low orbit satellite provides means to locate existing emergency locator beacons.

Anderson, R. E.↗

CANShield: Signal-based Intrusion Detection for Controller Area Networks

Modern vehicles rely on complex cyber-physical systems made up of hundreds of electronic control units (ECUs) connected through controller area network (CAN) buses. However, the CAN bus attack surface is increasing due to advanced features in automobiles, making it prone to injection attacks. The ordinary injection attacks disrupt the typical timing properties of the CAN data stream, and the rule-based intrusion detection systems (IDS) can easily detect them. However, advanced attackers can inject false data to the signal level, maintaining the regular pattern/frequency of the CAN messages. Such attacks can bypass the rule-based IDS or any anomaly-based IDS built on binary payload data. To make the vehicles robust against such intelligent attacks, we propose CANShield, a signal-based intrusion detection framework for the CAN bus that consists of three modules. A data preprocessing module handles the high-dimensional CAN data stream at the signal level and make them suitable for any machine learning model. A data analyzer module consists of multiple deep autoencoder networks, each analyzing the time series data from a different perspective. Finally, an attack detection module uses an ensemble method to make the final decision. Evaluation results on a standard signal-based dataset show the effectiveness of the CANShield in detecting five advanced attacks.

Shahriar, Md Hasan↗

A Machine Learning Approach to Improve Air Traffic Management Initiatives

Collaborating closely with commercial air carriers and related organizations, the Federal Aviation Administration(FAA) regulates air traffic and ensures the safety and efficiency of air operations. Air traffic controllers make strategic decisions, such as delaying, rerouting, or canceling flights, partly based on guidance provided by the FAA’s Air TrafficControl System Command Center (ATCSCC). The guidance includes, among other things, control measures known asTraffic Management Initiatives (TMIs) designed to enhance safety and improve operational efficiency. TMIs play a crucial role in managing the demand and capacity within the U.S. National Airspace System (NAS). Two major TMIs that are routinely used (primarily to mitigate the adverse effects of bad weather) are Ground Delay Programs (GDPs) andGround Stops (GSs). In a GDP, flights destined for airports facing thunderstorm activity experience delays at their origin airports. This proactive approach minimizes the risk of routing aircraft through hazardous weather conditions and also replaces (fuel burning) airborne delays with ground delays. In a GS, a temporary restriction is imposed on the departure or arrival of aircraft at a specific airport or within a designated airspace. Although other TMIs (e.g., miles-in-trail) are also implemented as part of (air) traffic flow management in the NAS, the focus of this work is on GDPs and GSs. Since TMIs, by design, lead to flight delays or cancellations, it is crucial to put in place the right set of parameters(e.g., scope and duration of the GDP). For example, when the end time of a GDP extends beyond what is necessary, it imposes unnecessary delays on departing flights. This situation could occur as a result of inaccurate prediction of the(required) duration of the GDP based on the weather forecast. On the other hand, if a GDP ends prematurely before the underlying capacity constraints are resolved at the destination airport, it may result in airborne holding. The delicate balance lies in matching the termination of the GDP precisely with the resolution of capacity constraints, avoiding both the imposition of unnecessary ground delays and the need for airborne holding due to premature program termination.Failing to specify the right parameters for TMIs also leads to flight delays, creating a significant obstacle in managing the increasing traffic volumes causing increased work load for the controllers. To address this issue, we propose the integration of Machine Learning (ML) models in the traffic flow management(TFM) pipeline. In current operations, decisions are made by human experts based on extensive training, historical patterns, available traffic and weather data. Since we have an abundance of data from past events that tell us the likely impact of various TMIs, by ingesting historical data, properly trained ML models can offer valuable insights and aid human decision-making. With the FAA increasingly exploring advanced analytics, ML emerges as a focal point for enhancing TFM within the National Airspace System (NAS). As a first step, this study aims to provide traffic controllers with decision-making support for the issuance and adjustment of TMIs. Data analytics and machine learning have been previously employed to address some of the challenges associated with TMIs. Numerous studies have concentrated on various facets of TMI issuance, exploring factors influencing TMI parameters, including arrival rate, airport capacity, and delay prediction. For example, using weather forecasts, several statistical methods were used to produce probabilistic capacity profiles which in conjunction with deterministic models provided insights into the GDP planning process [1–4]. The downside of using deterministic models is that they rely on fixed inputs and predetermined rules, which lack the ability to account for the inherent uncertainty and variability present in real-world scenarios. In a separate series of studies, researchers aimed to predict the occurrences of GDPs and GSs. The majority of these studies utilized various supervised learning methods, including Decision Trees, Naive Bayes, Support VectorMachines, and Random Forests to analyze the influence of weather conditions and arrival demand on TMI incidents[5–8]. However, these studies primarily focused on predicting the incidence of TMIs without explicitly addressing the scope of TMIs, including their duration and their geographical coverage. Furthermore, the emphasis of these studies was largely on GDPs, given their higher frequency and longer duration when compared to GSs. A limited number of studies focused on predicting the parameters of TMIs, specifically addressing their duration and extent. In one such study focusing on optimizing the TMI parameters at San Francisco International Airport (SFO),the authors utilized a probabilistic forecast of fog [9]. They simulated various capacity scenarios based on the (fog)burn-off forecasts, selecting GDP parameters that minimized airborne and overall ground delays. However, this approach exclusively emphasizes stratus (fog) burn-off as the primary determinant of GDP and GS, neglecting other influential factors like severe weather events, runway closures, lower capacity than traffic demand, and other important variables. Given the complexity of predicting the TMI and determining its scope, we seek a more holistic approach. We aim to consider all significant factors that could impact TMIs and their parameters. What sets this research apart is the fusion of all data sources relevant to the issuance and adjustment of TMIs and it represents the first comprehensive attempt to optimize TMIs in this manner. Since this comprehensive solution involves various aspects, we break down the problem into smaller components and input all parameters into a unified model called the “TMI Adjuster”. Figure 1 shows the overall framework and the list of datasets used in each model. The objective of the TMI Adjuster module is to deliver reliable, consistent and expedited recommendations for the progression, adjustment, and termination of TMIs. The ML solution entails developing a pipeline capable of predicting the necessity of a TMI (e.g., GS or GDP) along with its various parameters. For example, in the case of a GS, this includes the scope of the GS either in terms of distance from the destination airport or based on pre-defined airspace sectors. Here, scope refers to those regions and departing airports that are subject to the GS. In this paper, we concentrate on the issuance of GSs in the three major airports in the New York area — LaGuardia(LGA), John F. Kennedy International (JFK), and Newark Liberty International (EWR). We fuse traffic, weather and other relevant aviation data from years 2017 to 2019 to train and validate the ML models. In particular, we use the following datasets: •Terminal Aerodrome Forecast (TAF): meteorological forecasts specific to each airport, issued four times a day, covering predefined time periods. •TMI data: includes all GSs and GDPs along with their respective parameters. •Aviation System Performance Metrics (ASPM): includes traffic related data such as aircraft delays, arrival, and departure rates. •Notices to Airmen (NOTAMs): utilized to extract runway closure data and manage interdependencies between terminals in close proximity. •Flight cancellation data •Airspace Flow Programs (AFP): includes information on flight airborne holdings caused by TMIs. The data preprocessing entails transforming ASPM, TMI, AFP, NOTAMs, and weather data into an hourly format and consolidating all datasets by merging them based on date and time as the primary key. The TMI Adjuster framework comprises two parallel models: one dedicated to GS and a second model focused on GDP. As previously mentioned, our specific focus is on the GS model as a multi-classification problem. In this framework, each data point of the GS model input summarizes ten hours of data. Specifically, the data loader for the GS model generates the input and output of the model as follows: at a given time step, the input includes the actual traffic, weather, and TMI data from the two-hour window before the time step, alongside the weather forecast and scheduled traffic for the next 8 hours starting from the time step. Based on this information, the output of the GS model for each time interval consists of three dimensions. The first dimension represents a binary decision on whether there should be a GS in place for the next hour or not. The second dimension is related to the scope of the GS in the United States, and the third dimension is related to the scope of the GS in Canada (i.e., to determine if the GS impacts airports in Canada).One of the challenges with TMI modeling is the sparsity of TMI events, particularly regarding its scope. To address this challenge in the scope of the GS model output, we implement grouping. The GS scope for the US region is defined based on a list of centers that should be included when the GS is in place. With 20 centers in the US, we utilized historical data to group them into 4 categories. In particular, we summarized our historical data in a graph format where nodes represent centers, and link weights are defined based on the co-occurrence of centers in the scope parameter ofTMIs. By identified strongly connected components in this graph, we were able to partition the centers into four groups. We consider two model structures for the GS Model. Firstly, a hierarchical classification model [10], where the human decision-making for a GS is of hierarchical nature. The decision-maker first decides whether there is a need fora GS, and if the answer is yes, determines the scope. A hierarchical classification model organizes the problem into a class hierarchy, typically a tree or a Directed Acyclic Graph (DAG) structure, and considers the dependency of the decision in the previous step to the next component [10]. Here, we employ the local classifier per level approach, which involves training one multi-class classifier for each level of the class hierarchy. The second structure is the independent structure. In this setting, as the name suggests, we do not consider the dependency of the decisions in the different dimensions of the output of the model. Instead, for each dimension, we train a multi-class classifier independently. Table 1 summarizes GS model statistics for training, validation and testing. The table documents the effect of limiting data to the time steps when there was actually a TMI in place or when a TMI had just terminated. This resulted in a more balanced distribution of the GS class(GS positive class)versus “No GS”(GS negative class), which might help the training process. While JFK and LGA follow very similar distributions, with 40% and 42% GS positive class respectively, EWR has proportionally fewer GS incidents at 28%. Our subsequent phase involves evaluating the performance of both hierarchical structure and independent structure using different state-of-the-art multi-class classifier models such as Random Forest, Decision Trees, K-nearest Neighbors, and Logistic Regression and forecast the duration and scope of the GSs.

Farzan Masrour Shalmani↗