Agent evaluation

Task 08

Social media

The brief

Task 08 — Social media launch and responses

Use BRAND.md to create SOCIAL.md for launch day. Deliver:

  1. One LinkedIn post (120–180 words).
  2. One Bluesky post (maximum 280 characters).
  3. One Instagram caption (80–130 words, no more than four hashtags).
  4. Replies to all six sample comments. Each reply must fit the platform and be specific, calm, and honest.

Do not invent customer counts, performance claims, certifications, or roadmap commitments. Avoid engagement bait and canned “thanks for your feedback” prose. The outage comment needs acknowledgement and a useful next step without exposing private account details.

Write RESPONSE.md explaining the voice choices in at most 150 words. Work only in this directory.

Inputs given: BRAND.md

Scores

Criterion (max)Opus 5.5Sonnet 5.5Fable 5.1mimoGPT-6 AstraGPT-6.1 SolGPT-6 LunamuseGrok 4.7MiniMax M3.1 FlashMiniMax M3
voice/platform fit (3)2.752.752.752.752.52.252.252.251.7521.75
factual discipline (2)21.751.51.5221.751.251.250.50
response quality (3)2.752.752.752.52.252.2522.251.51.751
constraint adherence (2)22222222222
Total (10)9.59.2598.758.758.587.756.56.254.75
Grader's notes

Letters in the grader's text: A = Opus 5.5, B = Sonnet 5.5, C = GPT-6.1 Sol, D = GPT-6 Luna, E = MiniMax M3.1 Flash, F = Grok 4.7, G = GPT-6 Astra, H = mimo, I = Fable 5.1, J = MiniMax M3, K = muse.

G and H tie at 8.75. G is fully fact-clean but hedgier; H has the better voice but asserts a few things. G is ranked first on factual discipline. F (6.5) is ranked above E (6.25): E reads better but invents far more product and support behaviour. Every submission meets the numeric limits. E's RESPONSE is about 148-151 words depending on whether the header and dashes are counted, so it is treated as passing. Many submissions say flatly that papers are not used for training. That is a mild inference from local processing, penalized only lightly; asserting that a lab plan doesn't exist is penalized more.

Evaluation 9.5 / 10 graded blind as submission A

Clean, warm and precise copy that stays inside BRAND.md throughout. Replies that touch unknowns say 'haven't announced' or 'don't want to guess' rather than asserting anything exists or doesn't. The outage reply is well built. The only soft spot is the training reply, which reframes the question rather than answering it outright.

Strengths

  • 'Summaries with their receipts attached.' fits the mildly nerdy voice
  • Lab reply: 'we haven't announced a lab or institutional plan' plus a concrete DM next step
  • Outage reply apologises, moves to DM 'so your account details stay private', asks only for OS, file count, and screen state
  • Japanese reply proposes a specific trial test without claiming support

Weaknesses

  • Training reply ('Local means... aren't uploaded') never says yes or no to training; it offers specifics on request instead
  • LinkedIn chatbot reply is long (98 words)
Evidence the grader checked
  • LinkedIn 160 words
  • Bluesky 271 chars
  • Instagram 84 words incl. hashtags (81 without), 3 hashtags
  • RESPONSE ~146 words
  • 6/6 replies

Objective checks

  • RESPONSE.md ≈ 145 words (≤ 150)

Files

opus-5.5/08-social-media/RESPONSE.md

Voice choices

I kept it precise and warm by leading with what Threadmark does in plain terms, then giving the one concrete idea worth being mildly nerdy about: a summary you can check, "with its receipts attached." Every factual claim comes from BRAND.md: Windows and macOS, local processing, $12/month after a 14-day trial, and Linux as investigated but not promised. I added no customer counts, benchmarks, certifications, language support, institutional pricing or dates. Length fits each platform: LinkedIn explains, Bluesky compresses, and Instagram opens on a scene people will recognise. Replies answer the actual question first and say "I don't know" or "not yet" where that's the truth. Where more is needed, they offer one concrete next step, such as a DM, a trial test or specifics on request. The outage reply apologises, keeps account details out of public view and asks only for diagnostic facts.