Big Data · AI · Web

Maritime traffic analysis from AIS data

A third-year project in three parts — Big Data, artificial intelligence and web development — built on AIS data (vessel positions) from the Gulf of Mexico.

  • Third-year project
  • ISEN Nantes
  • 2024

The three parts

Each part is covered in seven steps, illustrated with figures from the project.

Big Data analysis

  1. Data collection and cleaning

    Acquisition of millions of AIS data points in the Gulf of Mexico. Handling of missing values and outliers, then removal of duplicates to guarantee the quality of the analysis.

    Validity rules applied to each variable
  2. Descriptive statistics

    Univariate analysis of the variables: latitude, longitude, speed (SOG), course (COG). Identification of navigation patterns and high-density maritime areas.

    Statistical summary of the variables
  3. Visualisations

    Charts to understand how vessels are distributed: by type (passenger, cargo, tanker), by speed and by length.

    Number of observations by vessel type
    Number of observations by cargo type
    Distribution of vessels by length
    Distribution of vessels by speed
  4. Finding the main ports

    Automatic extraction of the most frequent position pairs to infer high-density points, then verification on a map of the Gulf of Mexico.

    The busiest ports in the Gulf of Mexico
  5. Heatmaps

    Generation of heatmaps showing where maritime traffic is densest.

    Heatmap of maritime traffic
  6. Correlation analysis

    Study of the correlations between variables to prepare vessel type prediction: correlation matrix, statistical tests (Chi², ANOVA) and boxplots to validate the relationships.

    Correlation matrix
    Vessel length by type
    Mosaic plot: vessel type and status
  7. Flow diagrams

    Chord diagrams visualise the flows between the major ports, then the exchanges by vessel type, to analyse the preferred routes.

    Traffic flows between ports
    Passenger vessels (types 60–69)
    Cargo vessels (types 70–79)
    Tankers (types 80–89)

Artificial intelligence

  1. Trajectory clustering: first attempt

    A first, naive approach: K-means applied directly to the navigation data (COG, SOG, latitude, longitude…).

    Trajectories per cluster with naive clustering
  2. Comparing clustering methods

    Several methods tried on the same navigation data, followed by DBSCAN and agglomerative hierarchical clustering (AHC).

    Clusters obtained on raw data, first method
    Clusters obtained on raw data, second method
  3. Choosing the number of clusters

    Evaluation of the optimal number of clusters with the silhouette score, the elbow method, the Davies-Bouldin index (to minimise) and the Calinski-Harabasz index (to maximise).

    Four indices to choose the number of clusters
  4. Feature engineering and cluster visualisation

    Extraction of 10 descriptors per trajectory: geographic bounds, centres, extents, area covered, total distance. After normalisation and selection of the relevant parameters, K-means on these new features reveals distinct activity areas, consistent both spatially and by vessel type.

    Clusters obtained with the trajectory descriptors
    Trajectories coloured by cluster
  5. Predicting the vessel type

    A classification model predicts the vessel type; it is evaluated with a confusion matrix and cross-validation. An extra feature, the vessel's estimated capacity computed with a naval architecture formula, is justified by the correlation matrix and the boxplots.

    Results on the test set and with 5-fold cross-validation
    Correlation between the variables and the vessel type
    Added features by vessel type
  6. Predicting trajectories

    A MultiOutputRegressor predicts latitude and longitude 5, 10 and 15 minutes ahead. The predicted positions (red flags) follow the actual positions (blue dots) very closely.

    Actual and predicted trajectories
    Zoom on part of the route
  7. Validation on the real AIS database

    The whole approach was then run again on the actual maritime traffic data. Only the new clustering is shown here.

    Clusters obtained on the real AIS database
    Cluster composition

Web development

  1. Application architecture

    A complete web application (HTML, CSS, JavaScript, PHP, Python scripts, SQL) to explore the results of the analysis interactively.

    Application home page
  2. Vessel viewer

    Vessel search and display with a dual view, detailed list and interactive map. Buttons trigger the type and trajectory predictions for the selected vessel.

    Vessel list and interactive map
  3. Add-a-vessel form

    A detailed form to register new vessels, with their static information (MMSI, name, dimensions) and dynamic information (position, speed, course). Data is validated in real time and stored in an SQL database.

    Add-a-vessel form
  4. Interactive geospatial clustering

    Computation and display of vessel clusters on an interactive map, with dynamic filters (length, average speed, cargo type) to spot activity areas and navigation patterns.

    Vessel clusters on the map
  5. Vessel type prediction

    Three models (Random Forest, logistic regression, Gradient Boosting) classify the vessel automatically; each model's confidence score is displayed for comparison.

    Comparison of the three models for one vessel
  6. Trajectory prediction

    Prediction of a vessel's trajectory with a choice of horizon (5, 10 or 15 minutes). On the map, blue dots are actual positions and red dots are predicted ones.

    Actual and predicted trajectory
  7. Client-server interface

    Buttons test the CRUD requests (create, read, update, delete) of the REST API, and a log shows every operation performed on the database in real time.

    API responses and operation log