Goal-Oriented VQA Smoke Benchmark#
This dependency-light visual question answering path exercises the platform shape needed for goal-oriented semantic communication.
The reference semantic-packet smoke recipe is recipes/vqa_goal_oriented_smoke.yaml. It is a
plumbing example, not a canonical publication benchmark.
It runs this block sequence:
source.vqa_smoke
-> foundation.vqa_semantic_select
-> foundation.vqa_payload_encode
-> channel.bit_boundary
-> channel.identity_encoder
-> channel.bit_boundary
-> channel.identity_link
-> channel.bit_boundary
-> channel.bit_count_match
-> channel.identity_decoder
-> foundation.vqa_payload_decode
-> foundation.vqa_answer_from_packet
-> metrics.vqa
The source emits tiny built-in images, questions, ground-truth answers, and ground-truth boxes. The selector is intentionally local-rule based: it produces a vqa.semantic_packet.json with the question-relevant label and region. The payload codec serializes that semantic packet as JSON UTF-8 bytes and then as canonical unpacked uint8 bits.
This is not a final VQA model. BLIP, LLaVA, Qwen-VL, SAM, or CLIP adapters can replace
foundation.vqa_semantic_select and foundation.vqa_answer_from_packet while preserving:
typed VQA artifacts,
fixed channel bit boundaries,
payload and transmitted bit accounting,
VQA answer accuracy/exact-match metrics,
benchmark catalog validation.
Run it from the CLI:
uv run noema recipe run recipes/vqa_goal_oriented_smoke.yaml
uv run noema benchmark validate benchmarks/vqa_goal_oriented_smoke_v1.yaml
In the dashboard, open Template Recipes, expand Visual question answering, choose the desired starter, and press Run All. The Overview shows VQA as the recipe’s read-only purpose tag.
Pretrained VQA Baseline#
The recipe recipes/vqa_pretrained_transformers_smoke.yaml uses the registered operation foundation.vqa_transformers_answer, an optional full-image VQA answerer. It is useful as a pretrained task baseline, but it is not itself a communication codec and does not exercise the channel blocks. The default model is dandelin/vilt-b32-finetuned-vqa; Salesforce/blip-vqa-base is a practical alternative.
This adapter needs the foundation extra:
# Explicitly supplied wheel (Noema is not yet on PyPI):
python -m pip install \
"noema-lab[foundation] @ file:///absolute/path/to/noema_lab-<version>-<platform-tag>.whl"
# Source checkout:
uv sync --extra foundation
Use it when you want to compare the semantic-packet channel recipe against a conventional pretrained VQA model, or when designing a later VLM-based semantic encoder/decoder.