PoseImageNet

PoseImageNet

Pose Estimation for Extensive Classes Based on Rich Structure Prototypes

PoseImageNet is a pose dataset built on ImageNet, covering an unusually wide range of object classes and object structures. Because a single semantic class frequently contains objects that cannot be deformed into one another, the class is partitioned into structure prototypes — subsets of objects that share one keypoint definition. Every object pose in the dataset comes with its keypoint annotation and the prototype label that identifies the deformable set it belongs to.

01 — Definition

From a semantic class to structure prototypes

A semantic class is first split into subsets by structure, then pose is annotated inside each subset. The split is what makes pose estimation well posed for classes whose objects do not all share one skeleton.

Three structure prototypes from the Sunscreen semantic class, each represented by three deformable samples The Sunscreen semantic class belongs to the Toiletry superclass. Three structure prototypes occupy one panel each, and every panel contains three annotated samples arranged as a triangle with bidirectional arrows. Semantic class: Sunscreen Superclass: Toiletry Sunscreen · Model 116 keypoints · three deformable samples Sunscreen · Model 212 keypoints · three deformable samples Sunscreen · Model 316 keypoints · three deformable samples One semantic class is split into three structure prototypes; each prototype contains three mutually deformable samples. Three deformable samples for each representative structure prototype Three subfigures show a foldable phone, scissors, and excavator prototype. Each subfigure contains three independently drawn articulation states arranged as a triangle. Every state carries the same semantic skeleton for its prototype, and bidirectional arrows show that the three samples can deform into one another. Structure prototype examples three superclasses Device · foldable phonesame 4-point prototype · closed / half-open / openclosedhalf-openfully openmutually deformable Tool · scissorssame 6-point prototype · closed / half-open / openclosedhalf-openfully openmutually deformable Vehicle · excavatorsame 8-point prototype · boom low / mid / highboom lowboom midboom highmutually deformable same keypoint definition — TPS warp is defineddifferent prototype topology — no warp exists Representative structure prototypes across three superclasses Three subfigures show a device, vehicle, and plant prototype. Each subfigure contains three object samples arranged as a triangle. Every sample includes its own object silhouette and keypoint skeleton, and bidirectional arrows connect the three samples within a prototype. Structure prototype examples three superclasses Device · foldable phoneone 6-point structure prototypethree mutually deformable samples Vehicle · bicycleone 8-point structure prototypethree mutually deformable samples Plant · broadleafone 6-point structure prototypethree mutually deformable samples same keypoint definition — TPS warp is defineddifferent prototype topology — no warp exists
Figure 1. The Sunscreen semantic class from the Toiletry superclass is divided into three structure prototypes. Each panel shows one prototype through three annotated samples arranged as a triangle. Bidirectional arrows indicate that the samples share one keypoint definition and can deform into one another.
02 — Statistics

Scale and structure of the dataset

One overview of how the dataset is spread across superclasses, and two views of its structure: how many keypoints a prototype carries, and how many prototypes a semantic class splits into.

Prototypes per superclass

Distribution of prototypes by keypoint count

Each bar counts the structure prototypes with that number of keypoints.

Distribution of semantic classes by prototype count

Each bar counts the semantic classes containing that number of structure prototypes.

03 — Samples

Two prototypes sampled from each superclass

One semantic object class is selected from each of the 13 superclasses. Each row contains two structure prototypes from that class and three object poses from each prototype. Within each group, every pose is deformable into every other: same keypoint count, same keypoint correspondence, different pose.

Skeleton overlays are rendered from the supplied images and their matching keypoint annotation files.

04 — Annotation

How to read the annotation

Annotated acoustic guitar sample with 17 numbered keypoints and skeleton / 带有 17 个编号关键点和骨架的吉他标注样例
Figure 2. One annotated sample. Keypoints are ordered by semantic index within the prototype, so index i always denotes the same part across every image of that prototype.

Fields

key of annotationsAnnotation identifier (annotation_id), unique across the dataset.
annotations[id].image_idThe image this annotation belongs to (a key of images).
annotations[id].prototype_category_idThe structure prototype this sample belongs to (a key of category_net).
annotations[id].keypoint_xyFlat array [x₁,y₁, x₂,y₂, …] of original-image pixel coordinates; length = keypoint_number × 2.
annotations[id].keypoint_vVisibility per keypoint; length = keypoint_number, 0 = occluded, 1 = visible.
key of imagesImage identifier (image_id).
images[id].path_from_root_to_image_fileRelative path from the image root folder to the image file.
images[id].image_raw_width, images[id].image_raw_heightOriginal image width and height in pixels.
key of category_netStructure prototype identifier (prototype_category_id).
category_net[id].keypoint_numberNumber of keypoints defined for this prototype.
category_net[id].canonical_sample_idThe prototype's reference sample, given as an annotation_id.
category_net[id].skeletonConnected keypoint pairs, 0-based indices.
category_net[id].semantic_category_name, semantic_category_idName and identifier of the semantic class this prototype belongs to.
category_net[id].super_semantic_category_name, super_semantic_category_idName and identifier of the superclass.

Loading an annotation


        

Open download

A single ZIP archive bundles all annotation files (PoseImageNet.json, SampleTrack.json, and the per-prototype category definitions).

Download annotation ZIP Coming soon — the download link will be added once the archive is ready.
05 — SampleTrack

Restore the full dataset

SampleTrack.json maps each images[].id to its original image in ImageNet or UniKPT. The keys are strings; use sample_track[str(image_id)] to locate the source.

  1. Prepare the files

    Use matching versions of PoseImageNet.json and SampleTrack.json, plus the original ImageNet and UniKPT images.

  2. Locate each source image

    Look up the image ID. For ImageNet/images_train/…, resolve the remaining path under your ImageNet training folder. Resolve all other paths under the UniKPT images folder.

  3. Rebuild the image folders

    Copy every source image to output / images[].file_name and keep both JSON files. The restored images and annotations form the full dataset.

Example: image ID → original relative path

{
  "1": "ImageNet\\images_train\\n01682714\\n01682714_1007.JPEG",
  "117241": "300w\\images\\afw\\90800092_2.jpg"
}

Run the restoration

Download Python script

Save the script beside the two JSON files. Replace the source and output folders below, then run with Python 3. The script checks every source file before copying.

python restore_dataset.py --imagenet "D:/ImageNet/images_train" --unikpt "D:/UniKPT/images" --output "D:/PoseImageNet"

Images are copied at their original resolution. Image IDs, keypoint coordinates and skeleton connections stay unchanged.

06 — Citation

BibTeX


      

Related Open-Source Repositories