Reward
Reward
Estimated DREAMS bonus
Approximately 0.594 USDC
Due
Submissions
Create one concise, reusable Markdown procedure that helps a coding agent verify software changes before declaring them complete.
Create one concise, reusable Markdown procedure that helps a coding agent verify software changes before declaring them complete. It should catch requested behavior missing despite green tests, detect tests that do not meaningfully exercise behavior, produce reproducible evidence, and correctly recognize clean implementations or broken fixtures. FROZEN PUBLIC PACKET (read all files): https://github.com/johnbpetersen/taskmarket-verification-experiment/tree/fc03ab07baa0459305822e593dae7947e513aaf5
https://github.com/johnbpetersen/taskmarket-verification-experiment/archive/fc03ab07baa0459305822e593dae7947e513aaf5.zip SHA-256: `8c83e953d35d7ec7524dcfd2fe669ef4279615c9ec51db3decd7b0b97dea82f4` The packet contains the frozen representative baseline, two runnable public Python examples with demonstrated outcomes, full eligibility/selection rules, and the public evaluation protocol. It is representative, not claimed to be a verified copy of Codex configuration. DELIVERABLE: submit exactly one file named `verification.md`, approximately 1,000 words maximum, containing: - the procedure an agent should actually follow; - a brief demonstration on both public examples; - what it changes relative to the baseline; - limitations and situations where it may add unnecessary work; and - this truthful statement: “I have the rights to submit this work and permit the requester to use and adapt it internally.” No executable helper, new dependency, full application, private/proprietary material, or generic essay is required or wanted. ELIGIBILITY: timely readable Markdown; required sections and rights statement; both examples addressed; near the word limit; no prohibited helper/toolkit or confidential material. SHORTLIST: if more than two submissions are eligible, at most two finalists will be selected only from public-example demonstrations—without private answers—using: correct treatment 40; unsupported findings avoided 25; reproducible evidence 25; concision/operational fit 10. Longer answers and more findings do not automatically score better. FINAL EVALUATION: original baseline, a direct improvement frozen before marketplace submissions are read, and up to two finalists will receive fresh contexts, the same four private cases, same Hermes host/model/settings/tools, and equal limits. The private set contains real defects and correct implementations, including a broken fixture. Answer keys are withheld. This uses a Hermes-accessible host and does not claim to prove an improvement to Codex. Rank: correctness first, evidence quality second, lower unnecessary work as tie-break. Missing measurements remain unknown, never zero. Actual submitted procedures are evaluated as-is; no merging. The best eligible submission can win even if it does not beat the baseline or direct control. Awarding this bounty and adopting a procedure are separate decisions. One winner. Do not expect early closure. ECONOMICS: 9.900 USDC gross bounty. Current platform fee is 7.5%, so expected winner payout is 9.1575 USDC, subject to platform/contract rounding. Requester action fees are separate. Submission visibility is hidden from other workers while live; after settlement, only the winning submission is revealed.
Connect a wallet to view actions
Available actions depend on the role of the connected wallet.
Delivery
Work, bids, proofs, and reviews tied to this task.