Ani4D dataset (TBZ + Objaverse + Objaverse-XL, curated 2026-10-04; 16 cameras per clip since 2026-10-04)
The curated articulated-character dataset behind Ani4D (PoseAE codec + VoxelFM), plus the 50-asset Ani4D-vs-ActionMesh
benchmark under benchmark/. 9,975 kept assets: train 9,791 (TBZ 57 / Objaverse 1,935 / Objaverse-XL 7,799), val 134 (10 / 56 / 68), test = benchmark 50 (8 / 20 / 22); the split is poseae_split.json at the root and bench50.txt lists the
benchmark. Every kept asset is rendered from its front camera (9,975 of 9,975 transparent renders complete) and 5,166 of 9,975 so far from all 16 cameras (the 15 further views are still being rendered and published), every frame 1:1 with its pose sequence, is normalised to the frame-0
bounding box (centred, longest side 1), and passes the curation rules below. Statistics: dataset_stats_2026-10-04.md.
Layout, one folder per asset (records and ready-to-train cache merged):
inventory_<dataset>.jsonl per-asset scan record; bucket=excluded rows carry the curation reason
clean/<dataset>/<name>/
rig.json, motion.json bake records (skeleton + per-frame motion, source frame range)
frame_0000_anchor.glb rest-pose mesh (ground-truth topology); frame_0000_textured.glb, video*.mp4 where produced
manifest.npz rest mesh + skeleton + skinning weights
theta_motion.npz per-frame joint parameters (theta: rot 3 | stretch 1 | trans 3 per joint), frame-0 units
frames.npy per-frame ground-truth vertices [T, V, 3], frame-0 units
meta.json, RENORM_FRAME0.json, REALIGN_ROOT.json, TRIMMED.json provenance markers
slat_frame0.npz AniGen frame-0 structured latent (z_s, z_skin, z_skl on a 64^3 grid)
pose_vox_*.npz pose-voxel pools of the codec
frontview_rgba.tar THE conditioning render, camera 0: the levelled front view (az -90, el 0), 1024^2,
one straight-alpha RGBA PNG per pose frame (frame_%04d.png), uncompressed tar -- 9,975 assets, 668 GiB
multiview/view_01.mp4 .. view_15.mp4
cameras 1..15 (ActionMesh's 16-camera layout, camera 0 being the front view): camera k at
azimuth -90 + 22.5 k, elevation (20, 35, 5, 50)[k mod 4] degrees above the horizon, in the
asset's view frame; the front view's scene and framing, 1024^2, every pose frame. Each MP4
holds two LOSSLESS H.264 streams: stream 0 = the straight RGB (libx264rgb, -qp 0), stream 1 =
the alpha (gray, -qp 0), a keyframe every 16 frames -- 5,166 assets, 0.67 TiB
multiview/cameras.json the clip's cameras: layout, resolution, lens / sensor / field of view, framing (union-bbox
centre and diagonal, fill 0.8, headroom 1.15), view frame, source frame window, per view
azimuth / elevation / location / matrix_world (Blender world, Z up) / file / frames
frontview.tar the older OPAQUE render of camera 0 (the same scene over a grey world), kept because the
published models were trained on it; it equals the transparent render composited as below
video_feats_crop_g.npz DINOv2-giant features of the opaque render's subject crop, published 2026-09-21 where they
existed then; DERIVED and no longer maintained (the code computes features on the fly):
scripts/tools/recache_video_crop.py --res 518 --model dinov2_vitg14_reg rebuilds them
Reading the views: ffmpeg -i view_01.mp4 -map 0:v:0 -f rawvideo -pix_fmt rgb24 - gives the RGB frames and
-map 0:v:1 -pix_fmt gray the alpha, byte for byte the rendered PNGs (each MP4 was decoded and compared before upload).
The opaque render is the transparent one composited onto Blender's grey world: ground = 171 + d, where d is Blender's fixed
8-bit output dither (grey_render_dither_1024.png at the root: uint8 d + 1, the same in every frame, clip and channel);
opaque pixels are the RGB itself; a semi-transparent pixel is srgb(lin(rgb) a + lin(171) (1 - a)) + d (1 - a), blended in
linear light as Blender does. Measured on 50 clips, DINOv2-giant tokens of that composite sit 0.014 relative L2 from the
opaque render's (one frame of motion: 0.22).
The two PNG render directories ship as uncompressed tars, one per asset (tar xf frontview_rgba.tar):
unpacked they are 1.69 M of the dataset's 1.83 M files -- 92 % of the file count for 36 % of the
bytes -- and a HuggingFace repo is capped at 1,000,000 files. The 15 extra views are one MP4 per camera
(16 files per clip with cameras.json). Everything else is stored as ordinary files.
Curation (applied in this order; every removal is recorded in the inventories): texture present and shaded; T >= 32; no idle clips (fastest joint p95 < 1 deg/frame and root travel < 0.5 bone lengths); no implausible motion (> 90 deg/frame, jerk spikes, root jumps); theta->GT reconstruction <= 5 % of the bbox diagonal; grey/magenta placeholder renders; single-joint rigs; vertex count <= 300k; blank renders (subject < 0.1 % of the frame); render frame count == T; rest joints on the mesh; eight manual review rounds; a texture-patch audit (lost textures, see-through bodies; 2026-09-30); and a review of EVERY clip (2026-10-01): the maintainer's drop marks plus two Gemini passes -- a content review and a shading review (unlit / looks 2D) -- applied wherever the maintainer had not marked the clip Keep; and a consistency pass (2026-10-04): four drop marks made after the 10-01 application, a two-joint rig with a zero-length bone and a ten-vertex skinning test rig. Some renders carry explicit, per-asset repairs (a 3ds Max layered face decal, six TBZ animals' wing / hair card textures, a camera view frame for an asset whose source is not Z-up); the code repo's scripts/texture_repairs.json and scripts/render_views.json list them. Skeletons: one connected tree, no phantom leaves, root realigned onto the first skinned joint (UniMate's filters). Every benchmark skeleton is unseen in training: the TBZ part is two whole held-out animals (goat, seagull), and no training asset shares an Objaverse/XL benchmark asset's rig instance (joint set, hierarchy and rest bone lengths).
Sources & licensing -- derived from Truebones Zoo (TBZ), Objaverse and Objaverse-XL. The underlying assets keep their own licences; verify them before redistributing derived data.
- Downloads last month
- 1,175