M1 project · Computer vision

Assessing a person's cooperation for face recognition

Face recognition assumes that the person cooperates: uncovered face, gaze towards the camera, normal gait. This M1 project, carried out by two students over three months at ISEN Nantes, measures that degree of cooperation automatically from a plain video.

  • M1 project
  • ISEN Nantes
  • Team of two
  • 3 months
Three people tracked, each with a stable identity

Key figures

  • 4 + 2functional blocks and complementary modules
  • 3 monthsof work, as a team of two
  • 5sprints with an agile method
  • 0identity switches on the two-person test sequence

Context

The problem

Under controlled conditions, face recognition systems exceed 99% correct verification. In an access-control corridor or a public space, their performance drops as soon as the person does not cooperate: averted gaze, covered face, tilted head, running past.

The goal

Assess a person's degree of cooperation automatically, in real time and without disturbing them. The aim is not to replace face recognition, but to give it contextual analysis:

  • spot the behaviours that disturb it (profile view, hidden face, running, unusual posture);
  • produce a cooperation score usable by an operator or a downstream system;
  • keep the same identity for each person, even when they disappear for a moment.

The setting

M1 project at ISEN Nantes, lasting three months, carried out as a team of two: we worked together on every task. It was supervised by Khadidja Ould Amer and Cyril Barrelet.

Organisation

An agile method in five sprints, each delivering a feature that can be tested on its own: state of the art and architecture, face and gaze modules, gait analysis, complementary modules, optimisation and wrap-up. The code lived in a private GitHub repository, with one branch per module.

Constraints

Run on a standard processor, without a dedicated graphics card, accept a video file or a webcam as input, and keep no biometric data by default.

Architecture

Four independent blocks, each runnable on its own, written in Python.

Overall diagram of the system: tracking feeds four analyses in parallel, which converge towards the cooperation score

1 · Tracking and re-identification

Detect each person, track them from frame to frame and keep the same identity throughout the sequence.

2 · Gait and posture

Estimate the skeleton and walking speed, and spot running, unusual gait or a person on the floor.

3 · Face occlusion

Measure, zone by zone, what hides the face: hand, mask, object.

4 · Head pose and gaze

Estimate head orientation, gaze direction and eye opening to decide whether the person cooperates.

Tracking and re-identification

The central block: without a stable identity, no measurement can be followed over time.

  1. Detect and track

    YOLOv8m-seg detects people and provides a mask for each; BoT-SORT tracks them from frame to frame. The confidence threshold is lowered to 0.30 so as not to lose people seen in profile or half visible.

  2. Recognise by face and by body

    InsightFace provides a face embedding, kept only if the face is large enough, well detected and geometrically plausible, which rules out the backs of heads. A body embedding is computed on the person's mask. The two are fused, with the body weighing slightly more: clothing separates two to five people filmed together quite well.

  3. Decide without rushing

    A state machine waits eight frames before assigning an identity, keeps lost tracks so they can be re-associated, and tightens its thresholds when two people are close. A gallery keeps several prototypes per person, and a regular consolidation merges identities created twice.

    Re-identification state machine
    Three people cross paths: each keeps their number
  4. An offline colour-based variant

    For videos processed afterwards, a second pipeline measures torso colour on each person's mask, recovers the reference colours with K-means, then assigns identities frame by frame with the Hungarian algorithm. A diagnostic file flags ambiguous passages, which a corrections file can settle.

    The identity is recovered after leaving the frame
    Identity kept from behind as well as in profile

Gait and posture

YOLOv8-pose estimates the 17 skeleton keypoints of each person. Speed is computed from the displacement of the hip centre; the scale in metres is derived from the size of the skeleton, set against an average height of 1.70 m, and the measurement is then smoothed.

The classes

  • Still: below 0.3 m/s.
  • Normal walk: upright torso, symmetric arm swing.
  • Abnormal walk: torso leaning by more than 30° or arm asymmetry above 40%.
  • Running: above 2 m/s.
  • Lying down: torso close to horizontal.

A majority vote over ten frames keeps the class from flickering. The module was tested on the public Penn Action dataset and on our own videos.

Skeleton, speed and state of each person: one is still, the other is detected lying on the floor

Face analysis

Two blocks sharing the same face landmarks.

Face occlusion

MediaPipe FaceMesh places 478 landmarks on the face, which delimit anatomical zones: eyes, nose, mouth, cheeks, forehead. The person's skin colour is calibrated on the first frames, in the Y'CbCr space, which separates brightness from colour. A zone where too few pixels look like skin is considered covered.

The global score, from 0 to 100, weights the zones by their importance for recognition: eyes and nose count three times as much as the forehead. It was tested with a hand in front of the face, a medical mask, an object held up and varied lighting.

Head pose and gaze

Head orientation is estimated by solving PnP from six face points: nose, chin, eye corners and mouth corners. Eye opening is measured with the eye aspect ratio, and gaze direction from the position of the iris.

The decision

The person is judged cooperative if both eyes are open, if their gaze is on the camera and if their head is turned by no more than 25° sideways and 20° up or down. A vote over fourteen frames stabilises the display.

Complementary modules

Position within an authorised zone

The user draws the authorised zones on the image, for instance a pavement. They are converted into a mask, and at every frame the ground point of each tracked person is tested: a person spending more than 30% of the time outside the zone is flagged.

Best shot

For each identity, the algorithm keeps the most usable frame of the sequence. Its score combines face quality, how frontal the face is and the absence of occlusion.

People in green inside the authorised zone, in red outside it

Outcome

All six tasks of the subject were completed; the figures below come from the tests in the project report.

Re-identification

  • Simple sequence, 2 people: no identity switch.
  • Complex sequence, 4 people crossing paths: 2 switches.
  • Stress test, 5 people in similar clothing: 4 switches.

Frame rate

  • From 9 to 14 frames per second on CPU in our tests, depending on the number of people.
  • Models are exported to ONNX to speed up inference without a graphics card.

Observed limits

  • Beyond four people, real time is no longer guaranteed without a graphics card.
  • Skin calibration is sensitive to backlighting and strong lighting changes.
  • A face turned by more than 80° escapes the analysis.
  • The cooperation thresholds were tuned on few people.

Ways forward

  • A dedicated head-pose model, more accurate at large angles.
  • A lightweight network for occlusion, instead of colour analysis.
  • Tracking across several cameras.
  • A cooperation score learned over time by a temporal model.

Tools

YOLOv8 (Ultralytics)

Person detection and segmentation, skeleton estimation.

BoT-SORT

Multi-person tracking, chosen over DeepSORT for its better handling of occlusions.

InsightFace

Face embeddings (ArcFace) for re-identification.

MediaPipe FaceMesh

478 face landmarks, including the iris.

OpenCV

Video reading, colour spaces, PnP solving, annotated rendering.

ONNX Runtime · Google Colab

Optimised CPU inference; training on GPU.

Report

The specifications, the project management and the detail of each block.