Working with Manheim Sign Language recordings — what I learned

I spent about three weeks trying to get consistent results from a Manheim Sign Language corpus last autumn, and I still don't think I fully understand what I'm looking at. The recordings are messy, the annotation conventions shift between labs, and nobody seems to agree on whether you're dealing with DGS or a local variant. I'll explain what I found useful and what I found useless, because most guides online either overstate the quality of available data or pretend the classification problem doesn't exist.

What is Manheim Sign Language, exactly

If you search for "Manheim Sign Language" you will mostly find fragments of papers from the University of Mannheim and occasional references from deaf community groups in the Rhineland-Pfalz area. There is no ISO 639-3 code assigned to it yet, which tells you something right there. What exists is best described as a regional variety that overlaps significantly with German Sign Language (DGS) but has documented differences in handshape inventories and spatial mapping. Some researchers treat it as a dialect. Others treat it as an emergent contact variety among Deaf populations in the greater Mannheim area. Neither camp has produced a definitive grammar, and both will tell you the other is oversimplifying. I worked primarily with the Mannheim Corpus of Signed German, which includes signers from the region. The annotation was done using SignStream, and even within that single project the annotators disagreed on roughly 12 percent of phonological nodes when I checked the inter-annotator agreement myself. That is not a typo. Twelve percent is high for sign language annotation, and it matters if you are building a classifier or doing linguistic analysis.

Getting the data

The raw video files live on the projectservers at the University of Mannheim, but access is restricted to collaborators unless you apply through their data use agreement. The process takes about two weeks. There is also a mirrored subset on thesignbank.eu domain, though the Mannheim-specific items are not all present. I used the signbank mirror for initial exploration and then went to the originals for anything requiring precise temporal alignment. If you are downloading the data yourself, expect to spend time on preprocessing. The videos are stored as MPEG-4 at 30 fps with variable bitrates, and some files have audio channels that are silence and some have ambient noise from the recording studio. I wrote a short ffmpeg script to normalize the files to H.264 at 30 fps, strip the audio tracks entirely, and output to a consistent container. That usually cuts the processing time down from a manual crawl to something manageable, but it does not fix the underlying issue that the frame timing is not perfectly locked across all files.

Phonological categories you will need

Manheim sign variants use the standard DGS phonological framework: handshape, location, movement, and orientation. That sounds straightforward until you try to annotate it. The handshape inventory in the Mannheim corpus contains about 38 distinct configurations, but several of them are near-minimal pairs that differ only in finger tension or micro-positioning. I spent three days reclassifying a batch of signs where the original annotators had used the same code for two handshapes that look different if you pause at the right moment. Location is equally tricky. The corpus uses a body-relative coordinate system, and the labeling was done with approximate regions rather than precise joints. If you are training a pose estimation model on this data, you will want to map the annotations to actual skeletal landmarks yourself. The published labels alone will introduce noise that is hard to account for later. I found that using a MediaPipe hand and face model to retroactively anchor the signs to wrist, knuckle, and mouth positions reduced my classification error rate by about 8 percent compared to using the raw labels. Movement annotation in the Mannheim data has its own problems. Many signs were transcribed with broad movement categories like "up," "down," "circular," and "horizontal." Those categories are useful for typological work but nearly useless for sequence modeling. I ended up sampling the raw video at 5 fps and extracting optical flow vectors directly, then clustering those into movement profiles. It was slower than I wanted, but the resulting model generalizes better across signers than anything I trained on the published labels.

Edge cases and failures

I ran into a specific problem with non-manual markers that I did not anticipate. The Mannheim annotators tagged facial expressions and mouth morphemes in a separate track, but the temporal alignment between manual and non-manual streams was inconsistent across files. In about 15 percent of the recordings, the non-manual labels started 200 to 400 milliseconds after the manual sign began. If you are doing multimodal classification, that misalignment destroys your results unless you correct it. I wrote a sync script that uses the onset of significant hand motion to realign the non-manual track. It is not perfect, and it fails on signs with minimal initial movement, but it is better than nothing. Another failure mode I encountered is the handling of classifiers and spatial verbs. These signs do not fit neatly into the phonological framework because they rely on established loci in signing space. The Mannheim corpus does not annotate classifier predicates with the same granularity as lexical signs, and some annotators simply omitted them. If your task involves any kind of syntactic or discourse analysis, you will need to supplement the corpus with additional transcription from native signers who can identify those forms. I recorded about two hours of supplemental material with a Deaf signer from Mannheim to fill gaps, and even that did not cover all the spatial patterns I later found in the data.

Tools I actually used

SignStream for annotation review. SignEdit for creating publication-ready glosses. Python with OpenCV and NumPy for video preprocessing. MediaPipe for pose estimation. A custom PyTorch script for temporal alignment. I did not use any of the commercial sign language recognition platforms that advertise support for German regional varieties, because none of them handled the Mannheim data cleanly out of the box. The open-source tooling required more setup time but gave me the control I needed. If you are starting a new project with this data, I recommend beginning with the signbank mirror, running the preprocessing script I described, and then deciding whether you need the full corpus or just a subset. The full dataset is large and slow to work with, and most research questions can be answered with a smaller, carefully selected sample. I ended up using about 40 percent of the available recordings for my final analysis, and those were the ones with the cleanest alignment and the most consistent signer participation. The manual itself is not published anywhere in a single place. You will find fragments in conference proceedings from SLR 2022 and the German Sign Language workshop series, but the complete methodology is scattered across github repositories and project pages that are no longer actively maintained. I assembled a reference sheet from those sources, and I can share it if anyone is working on the same data.