Next after [BOORU CHARS 2025](https://nyaa.si/view/2004380) and [BOORU CHARS 2024](https://nyaa.si/view/1927862) pack with "the best of" new posts
from **danbooru** (safe and general, **ID 9100000..11200000 = 24.04.2025..20.04.2026**), **e621** (safe ONLY, **ID 5492000..6339620**)
and **zerochan ID 4430000..4680000** for the nearly same interval.
This time :
- image subsetting rely heavily on prefetched imageboard score, most of published images even not grabbed
- all Questionable images unconditionally placed to (upcoming) BOORU_ECCHI_2026 release
- it also will contain crossBOORU CATALOG updated since [2023](https://nyaa.si/view/1740396)
As usual :
- images initially filtered Mpixels>=0.48, shorter_side>=600 px, volume>=60000 bytes, no animations
stripes dropped or cropped to aspect ratio 0.4..2.1
- PNG/WEBP/AVIF converted to JPG using **cjpegli 96% quality** (2000000 bytes limit)
modest downsampling done to longer side 2560px (landscape) 1920px (1x1) 2480px (portrait)
- verbose file naming used **"%website% - %id% - %up_to_3_copyrights% ~ %up_to_5_characters% (%up_to_2_artists%).jpg"**
files uniquely identified by "%website%+%id%"
- some general image statistics got with EXIFTOOL and [IMAGE MAGICK](https://imagemagick.org)
- content analisys was mostly the same as BC2023-2025 with actual software and models
- [CRAFT text detector](https://github.com/fcakyon/craft-text-detector) used to estimate total size and number of text pieces
- torso components detected with [custom PyTorch model](https://github.com/aperveyev/booru_yolo/tree/main/models) being built over [Ultralitics YOLOv11](https://github.com/ultralytics/ultralytics)
- clustering and sorting inside cluster implemented to arrange compositionally and visually similar pictures
**inspect "readme" for details**
- images deduplicatied using [AntiDupl](https://github.com/ermig1979/AntiDupl) up to 2% similarity along with BOORU CHARS 2025, 2024
- semi-automated quality check and cleanup done as follow
- real-life photos, no-character landscapes, foods and macro thrown away
- most of comic and N-koma, overtexted images and line-arts filtered out
Beside images release contains tab separated texts :
- **BC_2026.tsv** file/image related metadata **700.010 rows**
- **BC_2026_tags.tsv** tags list with enrichment
- **BC_2026_yolo.tsv** detailed results for torso components detection
- **BC_2026_yolov11m_aa22.pt** PyTorch YOLOv11 model
and also additional "readme" with data descripton
Keep in mind this release (just like all others BOORU CHARS) is first of all
**a uniform dataset of character-centric art in effective local format suited for batch processing**
and then
**a representative catalog of anime/game/cartoon copyrights, characters and artists for visual estimation**
but
**not offer high image resolution and and does not claim to be complete.**
Comments - 0