Research internship · 4 months

Automated reading of capsule endoscopy videos

End-of-M1 internship at the Institute of Systems and Robotics of the University of Coimbra, Portugal. I designed a pipeline that goes through a capsule endoscopy video and reports each suspected lesion with its timecode, an annotated screenshot, its clinical category and its position in the digestive tract.

  • ISR · University of Coimbra
  • June – September 2026
  • M1 internship
A capsule films the digestive tract; a model analyses the video and locates each lesion

Key figures

All measured on videos of patients the models have never seen.

  • 71%of lesions found, at about one false alert per minute
  • 85%at the wider setting, about two false alerts per minute
  • 13 / 13patients with lesions flagged
  • 64 / 67lesions placed in the correct anatomical region
  • 0.90AUROC for bleeding on an external dataset
  • 21 minto analyse 48 min of video, without a graphics card

Context

The examination

Capsule endoscopy is the reference examination of the small bowel. The patient swallows a capsule the size of a large pill, which films the digestive tract for eight to twelve hours. The physician then reviews 50,000 to 100,000 frames: between thirty minutes and two hours of reading, during which a lesion visible on only a few frames can go unnoticed.

The goal

A pipeline that takes a complete video and outputs, for each suspected anomaly: its timecode, an annotated screenshot, its clinical category and its location — stomach, small bowel or colon, then position within the small bowel as a percentage of transit time. Knowing where a lesion lies matters as much as knowing what it is: the choice of intervention depends on it.

The setting

The internship took place at the Institute of Systems and Robotics (ISR) of the University of Coimbra, in the group of Professor Hélder Araújo, which works on computer vision and its medical applications. It was an individual research and development project, carried out in English in a highly international laboratory.

A reframed subject

The initial subject aimed at the capsule trajectory in centimetres. No public data can verify it: a single camera gives no scale, and the capsule advances, stops and moves backwards with peristalsis. I therefore proposed an anatomical localisation, which keeps the clinical purpose while remaining measurable.

Modest resources

Only public, anonymised datasets (Kvasir-Capsule, Capsule Vision 2024, WCEbleedGen), a laptop without a graphics card for video processing, and Kaggle's free GPU quota for training.

The pipeline

From raw video to report, in six steps.

Pipeline architecture: two branches share the same frames, one detects and names lesions, the other locates the capsule
  1. Video sampling

    The video is read with OpenCV at 2 frames per second: a 40-minute recording yields about 4,800 frames to analyse.

  2. Lesion detection

    An EfficientNet-B3 network fine-tuned on 12 classes: eight lesion types, normal mucosa, degraded views (bubbles, debris) and the two anatomical landmarks. The lesion probability of a frame is the sum of the eight lesion classes.

  3. Grouping into places to check

    Flagged frames less than two seconds apart form an episode, whose clinical category is voted by its frames; contiguous episodes form a zone. The physician thus gets one place to check, not fifty frames.

  4. Visual explanation

    Grad-CAM locates the image region that triggered the alert; it is boxed on the screenshot attached to each event.

  5. Anatomical localisation

    A second network, MobileNetV3-Small, recognises the pylorus and the ileocecal valve. These two crossings split the video into stomach, small bowel and colon; the position of a lesion within the small bowel is the time elapsed since the pylorus, divided by the total transit time.

    Position of three lesions as a percentage of small-bowel transit, on an 18-minute video
  6. Report

    One JSON file per video lists the events with all their fields, together with a timeline. Every model decision is saved frame by frame, so any rule can be replayed without re-running hours of computation.

An honest evaluation

Building the pipeline was not the hardest part: the real challenge was proving how good it is. A model evaluated on images of a patient it saw during training looks far better than it is: studies reporting over 95% mix images of the same patient, whereas with a split by patient the best published results on this dataset peak at a macro-F1 of 0.35.

  • Cross-validation by patient: four models, each trained without the videos of its group; every video is judged only by a model that has never seen it.
  • Held-out videos: five videos measured once, at the end, with thresholds frozen beforehand.
  • Criteria set before measuring: a change is adopted only if it meets a threshold decided in advance.
  • Data audit: a perceptual-fingerprint check, robust to cropping, verifies every external dataset before use. It removed 38,592 images duplicated between two public datasets.
  • Real pipeline versus measurement: the decisions of the complete pipeline are compared frame by frame with the measured ones, on all 16 videos.
Cross-validation by patient: each group is tested by the model that has not seen it
The main constraint: some lesions exist in only one or two patients

Results

Measured once on the held-out videos, with a tolerance of ± 5 seconds.

Lesions found on the 5 held-out videos: initial rule, final system, then wider setting
Correct detections of the final pipeline, each on a video never seen by its model: predicted category, Grad-CAM box and position

Detect

56 of 79 lesion episodes found (71%) at 1.2 false alerts per minute, against 66% for the initial rule at equal burden. At the wider setting: 67 of 79 (85%).

Name

53 of the 56 lesions found receive the right clinical category out of five: vascular, inflammatory, lymphangiectasia, foreign body, polyp.

Locate

64 of 67 lesions placed in the correct region, with a median error of 0.8 points on the position within the small bowel. A segmentation enforcing the anatomical order, measured separately, brings the pylorus error from 13 down to 5 seconds.

Bleeding

Almost all the fresh-blood images of the dataset come from a single patient: impossible to learn under cross-validation. Audited external images were added; on 2,151 never-seen images from another dataset, the AUROC reaches 0.90.

Clinical category given to lesion episodes: 127 of 164 correctly named (77%)
Probability of each section over a video: estimated crossings (red) and annotated ones (dashed)
External bleeding test on WCEbleedGen: AUROC of each model

Limits

A medical tool is judged as much by its limits as by its results. These are stated plainly in the report.

  • Three lesions in ten are missed at the default setting, while clinical software announces over 90% detection.
  • Too many alerts for clinical use: one alert per minute would mean hundreds of places to check on a full examination. They would need to be ranked by priority.
  • One camera, one hospital: nothing is validated on another capsule model.
  • Incomplete annotations: part of the “false alerts” are real lesions that were never annotated, which makes the figures pessimistic.

What comes next

Integrate the anatomical-order segmentation into the pipeline, validate on the 22 examinations of the Galar dataset filmed with the same camera in another hospital, and have a few videos fully annotated by a physician.

On a whole video, many events fall outside the annotated lesions: this is the main limitation

Tools

41 Python modules (about 6,200 lines), 10 training notebooks and over 60 commits.

PyTorch · torchvision

EfficientNet-B3 and MobileNetV3-Small networks, training, inference, ImageNet pre-trained weights.

PyTorch 2.9

OpenCV

Video decoding and sampling, drawing of annotated screenshots, homography trials.

OpenCV 4.13

Grad-CAM

Activation map locating the image region behind the decision.

scikit-learn

Section classifier, distance to normal mucosa (k-NN), AUROC and other metrics.

scikit-learn 1.6

NumPy · pandas

Caches of model outputs, replay of decision rules, result tables.

Kaggle (Tesla T4 GPU)

Training on the free quota: about 3 GPU hours for the final models, 13 minutes per model.

A method driven by measurement

The project moved in short iterations, each ending with a measurement. Several attractive ideas were dropped because they failed their criterion: this is what makes the final figures credible.

Adopted

  • Fine-tuning the detector under cross-validation by patient: more lesions found and fewer false alerts.
  • Grouped clinical categories rather than ten fine-grained classes that nobody names reliably from public data.
  • External bleeding data, after audit: no loss on the tuning videos.

Rejected, with the numbers

  • Generic foundation model DINOv2: AUROC of 0.70 against 0.92 for the in-house classifier.
  • Bubble detectors: they caused lesions to be lost.
  • Temporal smoothing of probabilities: no gain at equal detection.
  • Capsule progression by homography: impossible to validate without ground truth.
Internship timeline, from June to the end of September 2026

Report

The detail of every measurement, of the rejected options and of the evaluation protocol.

Endoscopic images: public Kvasir-Capsule dataset (CC BY 4.0 licence).