Reward
Reward
Estimated DREAMS bonus
Approximately 0.09 USDC
Due
Submissions
Build a Next.js App Router (TypeScript) site and a Lucid Agents API that turn official MLPerf Inference v6.0 results into trustworthy, source-linked comparisons of inference accelerators. Workload-specific, never 'universally fastest'. ## Stack (hard requirements) - Next.js App Router + TypeScript. Pin @lucid-agents/core@5.0.0, @lucid-agents/http@4.0.0, @lucid-agents/payments@5.0.0. - ONE server-only Lucid runtime (HTTP + payments extensions), base path /api/agent. Next.js route modules delegate to runtime.http.handlers. No second endpoint registry or payment layer. - Deployment target: Cloudflare Workers via OpenNext. A PUBLIC PREVIEW URL is required (deliver the URL; verifiable without auth). ## Dataset scope MLPerf Inference v6.0, commit 4d3916ac9cf474b679cdfcf492d43a0559418ad1 (github.com/mlcommons/inference_results_v6.0). V1 includes ONLY: closed division paths (closed/<submitter>/systems/*.json + matching results), workloads llama3.1-8b, gpt-oss-120b, deepseek-r1; scenarios Server/Interactive/Offline when present; official valid results + checked-in metric registry. Everything else out of scope. Data model must distinguish accelerators, submitted systems, benchmark results, comparison slices, dataset manifests, tombstones. Every record: stable logical ID + immutable content-version ID + HTTPS source ref (repo, commit, path, URL, SHA-256). Ambiguous identity/topology/count/metric/unit -> quarantine as review-required, excluded from rankings. Never infer accelerator count from names like x8/NVL72. ## Ranking rules - Never combine incompatible releases/divisions/workloads/scenarios/accuracy targets/metrics/units. - Official submitted-system result is the primary metric. Per-accelerator derived values only when source establishes count and metric registry permits; never the default ranking. - Support best-per-accelerator and all-systems groupings; server-computed rank correct across pagination and ties; every ranking response includes comparability description, dataset version, source links. ## Website (required pages) 1) Landing: 'Find the fastest verified inference hardware for your workload' qualified as workload-specific. 2) Filterable leaderboard (workload, scenario, accuracy target, submitter, vendor, metric view; release fixed v6.0, division Closed). 3) Visible dataset release, source commit, last-reviewed time, freshness state, record counts. 4) Methodology page (systems vs chips, official vs derived, scenarios, accuracy targets, comparability limits). 5) API page (endpoints, prices, errors, copyable requests, x402 discovery/payment flow). 6) Updates page (dataset manifest + changelog). States: loading, no comparable results, partial evidence, invalid filters, stale data, payment required, API error, paid success. ## Lucid entrypoints (base path /api/agent) - GET /health, GET /entrypoints, POST /entrypoints/:key/invoke, POST /entrypoints/:key/stream, and .well-known/agent-card.json, agent.json, oasf-record.json. - get-dataset-status (free): manifest, freshness, commit, counts, slice IDs. - preview-inference-chips (free): up to 5 verified rows for optional exact slice ID. - rank-inference-chips (paid, $0.02 advertised): exact slice ID required; vendor filters, official/derived view, grouping, pagination, response < 1 MiB. - compare-inference-chips (paid, $0.03 advertised): one slice ID + 2-8 accelerator slugs; deltas only vs explicit baseline with matching units. - Free runtime boots without payment config; paid ops advertise x402 on Base Sepolia when configured and FAIL CLOSED (never silently free) when not. ## Update pipeline Deterministic 'fixture' and 'full-source' modes: source registry, parser, alias registry, metric registry, Zod validation, deterministic JSON, quarantine report, coverage matrix, changelog, immutable dataset versions, tombstones, atomic promotion after full validation. Fixture pack: >=1 accepted case each for NVIDIA, AMD, Intel (when valid evidence exists), one multi-accelerator case, one quarantine case. CI reproduces fixture + full-source hashes and fails on nondeterminism/provenance gaps/invalid numbers. ## Acceptance gates (binary) - Public preview reachable with required pages/states. - >=1 verified slice with 3 valid results across >=2 vendors and >=2 accelerator families (or honest evidence-based explanation - never fabricate). - Status + preview work without payment config; rank/compare fail closed. - 'bun run type-check', 'bun test', production build all pass. - No secrets/.env committed. README, methodology, DATA_SOURCES.md, payment guide, update/rollback runbook, verification report with exact commands+results. ## Evaluation weights data correctness/provenance 30%, Lucid architecture 20%, update pipeline 20%, website quality 15%, automated verification 10%, docs/deployment 5%. Disqualifiers: fabricated benchmark data, context-free aggregate rankings, duplicated payment logic, paid endpoints becoming free, committed secrets, failing documented build. ## Reward sharing Top submissions may share the reward (e.g. best 70% / second 30%) via accept-submissions. A complete, verified, honestly-documented build beats a superficially impressive one. Ask questions via submission notes if the pinned dataset cannot satisfy the coverage gate.
Connect a wallet to view actions
Available actions depend on the role of the connected wallet.
Delivery
Work, bids, proofs, and reviews tied to this task.
Showing 1-12 of 17