Task-Oriented Benchmarks#
Task-oriented benchmark contracts support semantic-communication experiments where the main result is task success rather than pixel reconstruction.
Supported Task Contracts#
The research catalog now includes supported contracts for:
Task |
Metric operation |
Primary metrics |
|---|---|---|
Classification |
|
|
Visual question answering |
|
|
Object detection |
|
|
Segmentation |
|
|
Image-text retrieval |
|
|
Caption-to-image generation |
|
|
These operations are dependency-light evaluator blocks. The VQA evaluator uses one reference answer,
so it does not claim consensus-based standard VQA accuracy; its exact-match score is literal and
case-sensitive. The detection evaluator reports deterministic maximum-cardinality, label-matched
one-to-one precision/recall/F1 with both IoU and confidence thresholds encoded in each canonical
metric ID; it does not claim AP or mAP. Segmentation pools semantic-class rasters across samples and
does not claim instance-segmentation or COCO mask metrics. They define the task surface that heavier
model adapters can plug into later.
metrics.captioning remains available as a low-level artifact evaluator, but image captioning is
not advertised as a runnable task until a real captioning receiver is registered.
Artifact Boundary#
The task evaluators expect typed artifacts:
task.labels.json reference class labels
task.predictions.json predicted class labels
vqa.answers.json question-answer examples
vision.detections.json bounding boxes and labels
vision.segmentation_mask.numpy integer segmentation masks in NPZ
text.batch.json captions
retrieval.rankings.json ranked candidate ids plus bound candidate_ids/count/depth
retrieval.targets.json expected image id per text query, used by ranker blocks
image.batch.numpy generated image batches
foundation.embedding.numpy embedding matrices with model/revision/preprocessing-space provenance
VQA, detection, segmentation, CLIP retrieval, captioning, and generative receiver adapters use these stable artifact kinds even when their internal models differ.
CLIP Retrieval#
The repository includes a runnable image-text retrieval baseline:
uv run noema recipe validate recipes/clip_retrieval_smoke.yaml
uv run noema recipe run recipes/clip_retrieval_smoke.yaml
The recipe uses:
source.flickr8k_retrieval
-> foundation.clip_image_embed
-> foundation.clip_text_embed
-> foundation.clip_retrieval_rank
-> metrics.retrieval
The default source is real Flickr8k validation/test image-caption data, cached under
.noema/datasets/flickr8k, and the default backend is transformers_clip with
openai/clip-vit-base-patch32. Each query declares exactly one positive target, so the canonical
construct is Hit@K rather than multi-positive Recall@K. The ranker binds the complete candidate
universe, emitted ranking depth, and matching image/text embedding-space provenance; evaluators reject
duplicate or out-of-range K. Historical retrieval.recall_at_K outputs remain exact, explicitly
deprecated aliases only for this one-positive contract. source.retrieval_smoke remains only as a
fast synthetic debug source for plumbing tests; it
should not be reported as a research benchmark.
Detection And Segmentation#
The repository includes compact vision task baselines:
uv run noema recipe validate recipes/yolo_coco_detection.yaml
uv run noema recipe run recipes/yolo_coco_detection.yaml
uv run noema recipe validate recipes/yolo_coco_segmentation.yaml
uv run noema recipe run recipes/yolo_coco_segmentation.yaml
The detection recipe uses real COCO128 images and YOLO-format box labels:
source.coco128_detection
-> foundation.yolo_detect
-> metrics.detection
The segmentation recipe uses real COCO8-seg images and YOLO polygon labels:
source.coco8_segmentation
-> foundation.yolo_segment
-> metrics.segmentation
Both model adapters use pretrained Ultralytics YOLO models by default (yolo11n.pt and
yolo11n-seg.pt) and cache weights under .noema/checkpoints/yolo. In an installed environment use
an explicitly supplied wheel with
python -m pip install "noema-lab[vision] @ file:///absolute/path/to/noema_lab-<version>-<platform-tag>.whl";
in a source checkout use uv sync --extra vision
(or uv sync --all-extras) before launching these recipes.
Diffusion Generative Receiver#
The repository includes a generative receiver baseline:
uv run noema recipe validate recipes/diffusion_flickr8k_generation.yaml
uv run noema recipe run recipes/diffusion_flickr8k_generation.yaml
uv run noema benchmark validate benchmarks/diffusion_flickr8k_generation_v1.yaml
uv run noema benchmark run benchmarks/diffusion_flickr8k_generation_v1.yaml
The recipe uses:
source.flickr8k_retrieval
-> foundation.text_semantic_state_encode
-> foundation.semantic_state_payload_encode
-> channel.bit_boundary / identity channel / bit-count checks
-> foundation.semantic_state_payload_decode
-> foundation.diffusion_state_to_image
-> foundation.clip_text_embed + foundation.clip_image_embed
-> metrics.embedding_similarity
The default receiver is segmind/tiny-sd through Diffusers on CPU using the local-quality preset:
512x512 generation, 20 denoising steps, and a simple photographic prompt template. It is a trained
compact Stable Diffusion model, so the default produces an actual generated image rather than a
random test pipeline. It still downloads a model checkpoint, but it is far lighter than ordinary
Stable Diffusion releases. The recipe stores both the raw Flickr8k caption and the exact prompt
passed to Diffusers, generates an image from the received caption/SemanticState prompt, and reports
CLIP-space text-image alignment plus source-image/generated-image alignment. Install optional
dependencies before running:
# Explicitly supplied wheel (Noema is not yet on PyPI):
python -m pip install \
"noema-lab[foundation,retrieval-data] @ file:///absolute/path/to/noema_lab-<version>-<platform-tag>.whl"
# Source checkout:
uv sync --extra foundation --extra retrieval-data
This is a generative reconstruction benchmark, not a pixel-lossless reconstruction benchmark. The
source image is shown in the dashboard only as a visual and CLIP-space reference for the Flickr8k
caption pair. For a very fast plumbing-only smoke test, switch the UI quality preset to
Plumbing only; its tiny Diffusers pipeline output is not meaningful enough for research results.
Smoke Benchmark#
The first runnable pack is intentionally tiny:
uv run noema benchmark validate benchmarks/task_oriented_smoke_v1.yaml
uv run noema benchmark run benchmarks/task_oriented_smoke_v1.yaml
It uses source.task_labels_smoke to emit reference labels and candidate predictions, then evaluates
them with metrics.classification. This proves the task-success benchmark path, report generation,
catalog validation, and CLI result export.