Back to projects

AI HACK 2026 · Live

Relay

An AI agent that carries design requests from intake and clarification to delivery and approval in one place.

A design-request workspace built by a designer and me at AI HACK 2026 in Tokyo, whose theme was AI agents that make business work autonomous. I owned all development, including a Go worker that stores AI output only after matching it to verbatim source quotes, and a request flow that keeps working when the AI does not.

ResultLive service · public repo · 97 commits · evidence-checked AI output

Period
2026.09
Ownership
All development · Go API · River worker · AI integration · React
Team
2-person team (1 developer, 1 planner/designer)
Repository
Public repository

Handing the work around the creative work to an AI agent

My teammate, a designer, handled every clarification, prioritization, status update, and delivery message alone for requests from several marketers. For AI HACK 2026, themed around AI agents that make business work autonomous, we built a service that takes over that management work.

At the 2025 Korea-Japan Future Generation Forum, I argued that cooperation should move beyond slogans into hackathons and joint projects where young people actually build together. Relay is where I acted on that: competing at a hackathon in Japan and shipping a Japanese-language product for a Japanese creative workflow. We did not win a prize, but we finished and shipped it publicly.

How a request moves

Only validated AI output reaches the workflow

  1. 01

    Request form

    React · requester and manager views

  2. 02

    Go API

    Auth, permissions, request state

  3. 03

    Durable jobs

    PostgreSQL + River

  4. 04

    AI call

    OrcaRouter · bounded

  5. 05

    Evidence check

    Quote match · generation check

Store an AI summary only when its quote exists in the source

Request summaries let a designer act without rereading the original. If the AI added a deadline or condition that was never written, the summary would start the wrong work.

I made summaries extractive. Each item carries an evidence kind, ID, and quote, and the Go worker stores it only if the quote exists in the actual brief, answer, or comment and matches the item text.

The summary validator and the weekly-report validator are covered by unit tests. Unsupported output is never stored, and the report falls back to figures computed from records.

Constraints, implementation, and lessons

Telling the model to stick to facts is not a guarantee. The check had to live outside the model, in server code.

  • Requested JSON output and capped questions at 8 and summary items at 1–6.
  • Verified each quote exists at its source and matches the whitespace-normalized text.
  • Blocked delivery details as evidence for unfinished requests.
  • Rejected unknown fields and trailing JSON.

I learned to treat LLM output like any external API input: validate it before trusting it.

The requester's own words become the evidence for AI summaries

Separating AI status from business status

A finished AI analysis does not mean finished work. And while the AI runs, the requester may edit the request or the designer may move it forward.

I moved AI work into River jobs persisted in PostgreSQL and recorded a generation at dispatch. Before applying a result, the worker compares generations and discards it if the request changed. Only human approval can complete a request.

PostgreSQL integration tests cover a request completed during a provider call, duplicate and stale jobs, and preserving manually added tasks.

Constraints, implementation, and lessons

AI calls can take tens of seconds and must survive a closed browser tab, yet a late result must never overwrite what a person did in the meantime.

  • Accepted requests at submission and ran analysis in the background.
  • Discarded stale results by generation and recorded an analysis_stale event.
  • Rejected any AI attempt to mark work complete in code.
  • Capped attempts, output tokens, and execution time.

With async work, the first question is not when it finishes but whether its result is still valid when it does.

Completion is decided by human approval

Requests and deliveries keep moving without the AI

The theme was autonomous business work, but missing keys, rate limits, and model errors are certain in real use.

Every AI output is treated as a suggestion with a fallback: follow-up questions can be partially answered, delivery drafts fall back to editable templates, and weekly summaries fall back to computed figures.

Tests cover key encryption and masking, CSRF, and workspace isolation; the integration run with mocked OrcaRouter responses and the Docker build both pass.

Constraints, implementation, and lessons

Surfacing an AI failure as a blocking error would stop requesters from submitting and designers from delivering. Both the AI-assisted path and a human-only path had to remain.

  • Showed managers a settings link and requesters a notice when no key is configured.
  • Encrypted API keys at rest and excluded them from read APIs and model input.
  • Pinned the provider endpoint so keys are never sent to user-supplied URLs.
  • Enforced workspace and role permissions on the server.

An agent's reliability depends less on how often the AI succeeds than on whether people can act when it fails.

API keys are stored encrypted and never returned by the API

The service and GitHub repository are public. All 97 commits are mine, and the Go unit tests, PostgreSQL integration tests, frontend build, and Docker build pass.

For AI features, I now decide what happens when the AI is wrong and who acts when it stops before choosing prompts and models.

Not yet verified

We did not win a prize. Question quality, cost, and latency with live OrcaRouter models, email delivery, and time saved in real work have not been verified beyond mocked tests.

Next to verify

  • Measure missed information and unnecessary questions on real requests
  • Compare Japanese quality, JSON stability, and latency across models via OrcaRouter
  • Verify email delivery and deduplication in production