<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
    <title>Tianrun Qiu: Writing</title>
    <link href="https://r-q.name/blog/feed.xml" rel="self" />
    <link href="https://r-q.name/blog/" />
    <id>https://r-q.name/blog/</id>
    <updated>2026-10-07T12:00:00Z</updated>
    <author><name>Tianrun Qiu</name><uri>https://r-q.name/</uri></author>
    <entry>
        <title>Guardrails for an AI that publishes on its own</title>
        <link href="https://r-q.name/blog/fluentide-guardrails/" />
        <id>https://r-q.name/blog/fluentide-guardrails/</id>
        <published>2026-10-07T12:00:00Z</published>
        <updated>2026-10-07T12:00:00Z</updated>
        <summary>Fluentide's AI writes and publishes every episode, and every script has to pass its guardrails first. How we built them, from the misses that got past.</summary>
        <content type="html">&lt;p&gt;I co-built &lt;a href=&quot;https://fluentide.com/&quot;&gt;Fluentide&lt;/a&gt;, a Chinese listening app with 2,000+ learners. It turns &lt;a href=&quot;https://fluentide.com/library&quot;&gt;breaking news&lt;/a&gt; into Chinese episodes that sound native but stay at each learner&amp;#39;s level.&lt;/p&gt;
&lt;p&gt;Our AI has the autonomy to write and publish every episode. Before a script is voiced, it has to pass checks on level, repetition and new words, and AI reviewers flag any Chinese a native speaker wouldn&amp;#39;t say. I built many of those guardrails.&lt;/p&gt;
&lt;p&gt;This summer, an episode for Mandarin learners still went live in Cantonese, which they can&amp;#39;t understand, and none of our checks caught it.&lt;/p&gt;
&lt;figure&gt;
&lt;a href=&quot;https://r-q.name/blog/fluentide-guardrails/cover.png&quot;&gt;&lt;img loading=&quot;lazy&quot; decoding=&quot;async&quot; src=&quot;https://r-q.name/blog/fluentide-guardrails/cover.png&quot; width=&quot;2400&quot; height=&quot;1254&quot; alt=&quot;Two episodes our checks passed. The check said &amp;quot;Words at the right level&amp;quot;; a listener heard Cantonese, not Mandarin. The check said &amp;quot;New words reused enough&amp;quot;; a listener heard &amp;quot;It's a goose leg. It's a goose leg. It's not…&amp;quot;&quot;&gt;&lt;/a&gt;
&lt;/figure&gt;&lt;h2&gt;The episode in Cantonese&lt;/h2&gt;
&lt;p&gt;I noticed it myself. Some of our episodes came out in Mandarin and some in Cantonese, and I traced it to our text-to-speech service. It takes the language as a separate setting from the voice, and one part of our code never sent it. The service guessed. The script itself was fine, and every check we had only looked at the script.&lt;/p&gt;
&lt;p&gt;I locked the language setting and added a test that fails if it&amp;#39;s ever left out again.&lt;/p&gt;
&lt;h2&gt;Easy words, Chinese nobody says&lt;/h2&gt;
&lt;p&gt;We asked for words at the learner&amp;#39;s level. The AI kept to easy words but sometimes wrote phrases no native speaker would say, like 走海 (&amp;quot;to walk the sea&amp;quot;), or 有米饭吃 (&amp;quot;have cooked rice to eat&amp;quot;) where anyone would say 有饭吃 (&amp;quot;have food&amp;quot;). Our level check only counted vocabulary. It passed both, and they made it into a beginner episode before we caught them.&lt;/p&gt;
&lt;p&gt;They were easy to miss because of how our scripts are written. Each Chinese line sits next to its English translation, and the English tells you what the Chinese means, so you don&amp;#39;t notice when it&amp;#39;s wrong. A beginner would memorize a phrase like that as real Chinese. I built an AI reviewer that deletes the English and flags any line a native speaker wouldn&amp;#39;t say.&lt;/p&gt;
&lt;h2&gt;A goose leg on repeat&lt;/h2&gt;
&lt;p&gt;New words should come back a few times so they stick. Our level check had a floor on reuse but no ceiling. A draft for our lowest level took that literally and repeated whole sentences: &amp;quot;It&amp;#39;s a goose leg. It&amp;#39;s a goose leg. It&amp;#39;s not a goose leg…&amp;quot; It scored 99.8% known words and passed.&lt;/p&gt;
&lt;p&gt;I added a ceiling on repetition and set it by measuring all our published scripts. It sits high enough that normal scripts pass, and only real padding gets flagged.&lt;/p&gt;
&lt;h2&gt;The exception the AI leaned on&lt;/h2&gt;
&lt;p&gt;We let the AI leave a word out of the new-word count if the episode explained that word. It explained so many that one beginner episode hit about 15 new words a minute, against our usual 9.&lt;/p&gt;
&lt;p&gt;So we added a cap that counts every new word, explained or not. Over the cap, the AI is told to cut the words it uses only once first.&lt;/p&gt;
&lt;h2&gt;City guides with no facts&lt;/h2&gt;
&lt;p&gt;We also wanted almost every word to be one the learner already knows. The first batch of beginner city guides hit up to 99% by leaving out names and places. Half of them had no checkable facts at all.&lt;/p&gt;
&lt;p&gt;Now a check counts facts like years, distances and prices, with the bar set at our own news episodes. Numbers are some of the first words beginners learn, so the facts came back to our &lt;a href=&quot;https://fluentide.com/chinese/explore-china&quot;&gt;city guides&lt;/a&gt; without making them harder.&lt;/p&gt;
&lt;h2&gt;Testing the checks on real HSK papers&lt;/h2&gt;
&lt;p&gt;A check can also be too strict and block good work. That&amp;#39;s harder to spot, because a blocked script never reaches anyone to read. So I started testing our checks on work we knew was good. When our AI began writing &lt;a href=&quot;https://fluentide.com/tools/hsk-mock-test&quot;&gt;practice exams for HSK&lt;/a&gt;, China&amp;#39;s official Chinese proficiency test, every check ran on the official sample papers first. If a real exam question failed one of our checks, we fixed the check.&lt;/p&gt;
&lt;p&gt;Our plan said the right answer is almost always a paraphrase of what you hear, so we were going to reject questions where it&amp;#39;s said word for word. On the real HSK 4 paper, 9 of the 32 answers are said word for word. The rule would have failed 28% of the real exam. On the official papers, not one of 256 wrong answers is said word for word, so the check now rejects any question where a wrong answer is said word for word.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Whatever AI I work on next, the first thing I&amp;#39;ll look for is its goose leg.&lt;/p&gt;
</content>
    </entry>
    <entry>
        <title>Shipping the same day a runner asks: how I build with Claude Code</title>
        <link href="https://r-q.name/blog/claude-code-loop/" />
        <id>https://r-q.name/blog/claude-code-loop/</id>
        <published>2026-10-04T12:00:00Z</published>
        <updated>2026-10-04T12:00:00Z</updated>
        <summary>My loop with Claude Code on two AI products: sessions that message each other, research that argues against itself, and testing done before I look.</summary>
        <content type="html">&lt;p&gt;My Claude Code sessions message each other. They test my changes before I look, and they re-film my demo videos when the app changes. My median message to them is under 20 words.&lt;/p&gt;
&lt;p&gt;I spent the summer as an AI-native builder on two products: &lt;a href=&quot;https://fluentide.com/&quot;&gt;Fluentide&lt;/a&gt;, an AI-powered Chinese listening app with 2,000+ learners, and &lt;a href=&quot;https://taosport.cn/&quot;&gt;TAOSport&lt;/a&gt;, an AI running coach with 250+ active runners. Twenty words are enough because Claude runs a loop that starts and ends with the people using the app. Here&amp;#39;s that loop.&lt;/p&gt;
&lt;figure&gt;
&lt;a href=&quot;https://r-q.name/blog/claude-code-loop/loop.png&quot;&gt;&lt;img loading=&quot;lazy&quot; decoding=&quot;async&quot; src=&quot;https://r-q.name/blog/claude-code-loop/loop.png&quot; width=&quot;2400&quot; height=&quot;1254&quot; alt=&quot;A loop that starts and ends with users: Hear, Decide, Build, Check, and Back to the user, around Fluentide (2,000+ learners) and TAOSport (250+ active runners).&quot;&gt;&lt;/a&gt;
&lt;/figure&gt;&lt;h2&gt;Hear&lt;/h2&gt;
&lt;p&gt;Requests from people reach me in two ways. After a team meeting I paste the raw transcript: four agents read the code in parallel, and Claude comes back with a plan for every request in it. Bugs come from runners, as screenshots in our group chat, and a screenshot plus a few words is enough for Claude to reproduce the problem on its own.&lt;/p&gt;
&lt;p&gt;On Fluentide, though, more of the work started from our data than from anyone&amp;#39;s message. That loop starts in a different place, so it gets its own post, the next in this series. This one follows the requests people bring.&lt;/p&gt;
&lt;h2&gt;Decide&lt;/h2&gt;
&lt;p&gt;A plan is only as good as the research behind it, and Claude&amp;#39;s research comes back confident and well cited. So I make it argue against itself: it studies how others solve the problem, tries to prove each finding wrong with our own code, history and usage, and only then advises.&lt;/p&gt;
&lt;p&gt;That step often changes the plan. When Claude studied Telegram&amp;#39;s code and 20 other chat apps for our coach&amp;#39;s chat, it named our biggest problem, citing about ten past bug fixes. Made to prove itself wrong against our own code, it recounted: one. Six of its recommendations, including the one it ranked first, were overturned or scaled back before the plan was written.&lt;/p&gt;
&lt;p&gt;Most decisions are smaller than that, and they shouldn&amp;#39;t wait for me. Agents like to end a task with a list of open questions; mine settle the small calls with evidence and write the call and its reason into the commit. Only four kinds come back to me: logins, money, taste calls I keep for myself, and anything public or irreversible.&lt;/p&gt;
&lt;h2&gt;Build&lt;/h2&gt;
&lt;p&gt;Claude Code sessions can message each other. It&amp;#39;s called &lt;a href=&quot;https://code.claude.com/docs/en/cross-session-messaging&quot;&gt;cross-session messaging&lt;/a&gt;, and many people don&amp;#39;t know it exists. So once a plan is settled, I give each step its own session and let them coordinate. On one plan they agreed who owned which files, confirmed &amp;quot;226 tests pass on the merged file&amp;quot;, and passed findings across. I didn&amp;#39;t relay a single message. The most useful bug report came from the session next door: it caught our AI grader failing correct translations, and told the session testing the grader.&lt;/p&gt;
&lt;figure&gt;
&lt;a href=&quot;https://r-q.name/blog/claude-code-loop/sessions.png&quot;&gt;&lt;img loading=&quot;lazy&quot; decoding=&quot;async&quot; src=&quot;https://r-q.name/blog/claude-code-loop/sessions.png&quot; width=&quot;2400&quot; height=&quot;2118&quot; alt=&quot;Three real messages between Claude Code sessions on one Fluentide plan: one confirms 226 tests pass on a merged file, one proposes which files each session owns, and two minutes later the other confirms the split and reports that the AI grader's verdicts argue themselves out of the flag.&quot;&gt;&lt;/a&gt;
&lt;/figure&gt;&lt;p&gt;When an app changes, its demo video usually goes stale. Mine don&amp;#39;t, because Claude is also my video editor: it drives the real app on camera and times every cut from the measured narration, never by hand. One command re-films the demo and re-cuts the edit in about two minutes. One demo survived four rewrites of the app it was filming without me touching a single timing; its captions still light up word by word, and a click still lands exactly on the spoken word &amp;quot;back&amp;quot;.&lt;/p&gt;
&lt;h2&gt;Check&lt;/h2&gt;
&lt;p&gt;A change isn&amp;#39;t done when the code is written, but for most changes I don&amp;#39;t click through anything myself. Claude runs the change end to end in &lt;a href=&quot;https://lite.ego.app/&quot;&gt;Ego Lite&lt;/a&gt;, a browser built for agents, logged in like a real user, and fixes what it finds before reporting. Then it hands me a storyboard: one captioned screenshot per step, loading, empty and error states included. Scrolling through it is my review.&lt;/p&gt;
&lt;figure&gt;
&lt;a href=&quot;https://r-q.name/blog/claude-code-loop/storyboard.webp&quot;&gt;&lt;img loading=&quot;lazy&quot; decoding=&quot;async&quot; src=&quot;https://r-q.name/blog/claude-code-loop/storyboard.webp&quot; width=&quot;2400&quot; height=&quot;3794&quot; alt=&quot;Six frames of a storyboard from TAOSport's AI chat, translated from Chinese: an unanswered question with Retry, a failed retry, the answer, thumbs-down reasons, a table streaming in, and the switch-chat dialog.&quot;&gt;&lt;/a&gt;
&lt;/figure&gt;&lt;p&gt;Changes to our AI coach need one more check, since a screen can look right while the coach&amp;#39;s answer is wrong. Claude replays the exact words the runner complained about in the real app, saves the answer with screenshots and a full trace, then reads it and signs a verdict. The fix counts only when that runner&amp;#39;s own question gets a good answer.&lt;/p&gt;
&lt;h2&gt;Back to the user&lt;/h2&gt;
&lt;p&gt;That runner is also where the loop ends. It closes in the chat where it started: every change gets posted back to the group, and many requests shipped the same day a runner asked.&lt;/p&gt;
&lt;p&gt;That is why my messages can stay short. Claude does most of the building and checking. My time goes to the people in that chat, and to the calls only I can make.&lt;/p&gt;
</content>
    </entry>
    <entry>
        <title>I built an on-device AI that judges every screen in a second</title>
        <link href="https://r-q.name/blog/qualm/" />
        <id>https://r-q.name/blog/qualm/</id>
        <published>2026-09-25T12:00:00Z</published>
        <updated>2026-09-25T12:00:00Z</updated>
        <summary>My thesis app needed the cloud; on-device, one screen took 20 to 40 seconds. A new kind of model cut that to one second, and I built Qualm for the Mac in a day.</summary>
        <content type="html">&lt;figure class=&quot;post-video&quot;&gt;
&lt;video controls playsinline preload=&quot;none&quot; poster=&quot;https://r-q.name/blog/qualm/cover.jpg&quot; width=&quot;1280&quot; height=&quot;720&quot; aria-label=&quot;Qualm, a 61-second narrated video with captions&quot;&gt;
&lt;source src=&quot;https://r-q.name/blog/qualm/qualm.mp4&quot; type=&quot;video/mp4&quot;&gt;
&lt;/video&gt;
&lt;/figure&gt;&lt;p&gt;Almost no screen-time tool uses AI. The idea is obvious: judge every screen instead of blocking a site. But that is a model call every few seconds on everything you read. In the cloud it costs money and privacy; on the device it was too slow.&lt;/p&gt;
&lt;p&gt;I built the cloud version anyway. &lt;a href=&quot;https://seenot.site/&quot;&gt;SeeNot&lt;/a&gt;, for Android, was, as far as I know, the first to judge the screen itself, and it became my thesis at &lt;a href=&quot;https://www.sustech.edu.cn/en/&quot;&gt;SUSTech&lt;/a&gt;. It worked, and it never took off: people liked the idea, few downloaded it, and I never liked where the screenshots went.&lt;/p&gt;
&lt;p&gt;On September 15, &lt;a href=&quot;https://typesafe.ai/&quot;&gt;TypeSafe&lt;/a&gt; released &lt;a href=&quot;https://typesafe.ai/blog/introducing-system-one-models-and-jev&quot;&gt;Jev&lt;/a&gt;, a model that doesn&amp;#39;t write text. You give it a text and a list of questions with fixed answers (yes or no, or one option from a list), and it answers all of them at once, in well under a second; because nothing is generated, ten questions take about as long as one. That is exactly what my app had been asking a VLM to do with a page-long prompt. And &lt;a href=&quot;https://github.com/jaredpalmer/kev&quot;&gt;Kev&lt;/a&gt;, an open-source version of the same idea, does it on an Apple silicon Mac in about a second.&lt;/p&gt;
&lt;p&gt;What was new was where it could run. The on-device model I had tried for SeeNot took 20 to 40 seconds per judgment; this takes about one, with nothing sent anywhere. So I built the second try, Qualm, a Mac menu bar app.&lt;/p&gt;
&lt;p&gt;YouTube is a lecture and a Shorts feed; the lecture stays, and the Shorts get a pop-up that says why. Each new screen gets a short list of questions (what kind of page is this, what is it for, is it private, does it break each of my rules), and plain code decides. I no longer write a prompt; I choose the questions and write the code that acts on the answers.&lt;/p&gt;
&lt;p&gt;Most of what the first version taught me is a list of what the tool must never do. It never hard-blocks. It runs local by default, and you can switch to TypeSafe&amp;#39;s hosted Jev, but it never falls back to it quietly, because the tool reads everything you read. And rules are generalizable sentences, so &amp;quot;short videos made for endless swiping&amp;quot; catches a site nobody listed.&lt;/p&gt;
&lt;p&gt;A week after the Jev launch, I sat down with it, ran trials the first day, had a working app that evening, and have used it daily since. I built it with Claude Code, and for agents too: every setting is also a command, so an agent can add a rule, test it on your own recent screens, and undo it. The night before shipping I asked for an audit as a fresh user and went to sleep; by morning 77 subagents had filed 204 problems, two of them serious.&lt;/p&gt;
&lt;p&gt;You can try it at &lt;a href=&quot;https://qualm.r-q.name/&quot;&gt;qualm.r-q.name&lt;/a&gt;, and the code is &lt;a href=&quot;https://github.com/RoderickQiu/qualm&quot;&gt;on GitHub&lt;/a&gt;.&lt;/p&gt;
</content>
    </entry>
</feed>
