<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Better Harness Blog</title>
        <link>https://qoderai.github.io/better-harness/blog</link>
        <description>Better Harness Blog</description>
        <lastBuildDate>Fri, 14 Aug 2026 02:00:00 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <item>
            <title><![CDATA[Harness Inspector: Seeing an Agent Delivery from Intent to Commit]]></title>
            <link>https://qoderai.github.io/better-harness/blog/harness-inspector</link>
            <guid>https://qoderai.github.io/better-harness/blog/harness-inspector</guid>
            <pubDate>Fri, 14 Aug 2026 02:00:00 GMT</pubDate>
            <description><![CDATA[A session log tells you what an agent did, not why the task started or which actions reached the codebase. Harness Inspector reconnects intent, sessions, file activity, and Git commits into one traceable delivery you can actually read.]]></description>
            <content:encoded><![CDATA[<p>Lately we have been trying to improve one capability in Better Harness: <strong>automatic SKILL distillation</strong> — recognizing the recurring work paths inside an agent's real sessions, and then deciding which of them are worth capturing as a reusable SKILL. Once we actually started, we found the problem was far harder than "just analyze a session."</p>
<p>For a real software task, an agent's behavior never happens in isolation. It starts from a requirement or a user story, moves through understanding the intent, exploring context, editing code, and verifying the change, and only then produces a contribution someone can review. Looking at the session alone, you can see <em>what</em> the agent did, but it is hard to tell <em>why</em> those actions happened, or <em>which</em> of them actually made it into the final delivery.</p>
<p>So we began treating a single agent delivery as one continuous chain. Today, you only need to run this inside a project directory:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">npx @qoder-ai/better-harness inspector</span><br></div></code></pre></div></div>
<p>and you get a local, read-only Harness Inspector page that puts the project's agent sessions, file activity, and Git commits into one interactive view.</p>
<p>GitHub: <a href="https://github.com/QoderAI/better-harness" target="_blank" rel="noopener noreferrer" class="">https://github.com/QoderAI/better-harness</a></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="from-a-session-to-a-complete-delivery">From a session to a complete delivery<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#from-a-session-to-a-complete-delivery" class="hash-link" aria-label="Direct link to From a session to a complete delivery" title="Direct link to From a session to a complete delivery" translate="no">​</a></h2>
<p>Harness Inspector started as a session-debugging tool: a way to see what an agent said, which tools it called, and which files it changed. But as sessions, file activity, and Git history got wired together, we realized the thing worth observing was not the conversation itself, but this:</p>
<blockquote>
<p>How a software change starts from an intent, passes through an agent's execution, and ends up as an artifact that can enter the engineering system.</p>
</blockquote>
<p>The session is only the middle of that chain.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-delivery-chain-from-intent-to-output">The delivery chain: from intent to output<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#the-delivery-chain-from-intent-to-output" class="hash-link" aria-label="Direct link to The delivery chain: from intent to output" title="Direct link to The delivery chain: from intent to output" translate="no">​</a></h3>
<p>We split a coding-agent delivery into three continuous parts with different boundaries:</p>
<p><img decoding="async" loading="lazy" alt="Intent, process, and output form one traceable delivery chain from requirement to commit" src="data:image/svg+xml;base64,PD94bWwgdmVyc2lvbj0iMS4wIiBlbmNvZGluZz0iVVRGLTgiPz4KPHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHdpZHRoPSIxODAwIiBoZWlnaHQ9IjEwNDAiIHZpZXdCb3g9IjAgMCAxODAwIDEwNDAiIHJvbGU9ImltZyIgYXJpYS1sYWJlbGxlZGJ5PSJ0aXRsZSBkZXNjIj4KICA8dGl0bGUgaWQ9InRpdGxlIj5IYXJuZXNzIEluc3BlY3RvciBhcmNoaXRlY3R1cmU8L3RpdGxlPgogIDxkZXNjIGlkPSJkZXNjIj5BIGxvY2FsLCByZWFkLW9ubHkgZXZpZGVuY2UgcGlwZWxpbmUuIEZlYXR1cmUgVHJlZSwgY29kaW5nLWFnZW50IHNlc3Npb25zLCBHaXQgaGlzdG9yeSwgYW5kIEVudGlyZSBjaGVja3BvaW50cyBlbnRlciBhIGJvdW5kZWQgY29sbGVjdG9yLCBhcmUgcHJpdmFjeS1maWx0ZXJlZCBhbmQgbm9ybWFsaXplZCwgY29ycmVsYXRlZCB3aXRoIGV4cGxpY2l0IGV2aWRlbmNlIGxpbWl0cywgcHJvamVjdGVkIGludG8gSGFybmVzc0luc3BlY3RvclJlcG9ydFYxLCBhbmQgcmVuZGVyZWQgYXMgYSBzZWxmLWNvbnRhaW5lZCBIVE1MIHdvcmtiZW5jaC48L2Rlc2M+CgogIDxkZWZzPgogICAgPHN0eWxlPgogICAgICB0ZXh0IHsgZm9udC1mYW1pbHk6IEludGVyLCB1aS1zYW5zLXNlcmlmLCBzeXN0ZW0tdWksIC1hcHBsZS1zeXN0ZW0sIEJsaW5rTWFjU3lzdGVtRm9udCwgIlNlZ29lIFVJIiwgc2Fucy1zZXJpZjsgfQogICAgICAuc2VjdGlvbiB7IGZpbGw6ICM4ZTlhOTQ7IGZvbnQtc2l6ZTogMThweDsgZm9udC13ZWlnaHQ6IDcwMDsgbGV0dGVyLXNwYWNpbmc6IDJweDsgfQogICAgICAubGFiZWwgeyBmaWxsOiAjZjRmN2Y1OyBmb250LXNpemU6IDI2cHg7IGZvbnQtd2VpZ2h0OiA3MDA7IH0KICAgICAgLmxhYmVsLXNtIHsgZmlsbDogI2Y0ZjdmNTsgZm9udC1zaXplOiAyMXB4OyBmb250LXdlaWdodDogNzAwOyB9CiAgICAgIC5oaW50IHsgZmlsbDogI2FlYjliMzsgZm9udC1zaXplOiAxN3B4OyBmb250LXdlaWdodDogNDAwOyB9CiAgICAgIC5zb3VyY2UgeyBmaWxsOiAjMTcxYjFkOyBzdHJva2U6ICM1OTY1NWY7IHN0cm9rZS13aWR0aDogMjsgfQogICAgICAucHJvY2VzcyB7IGZpbGw6ICMxNTFmMjc7IHN0cm9rZTogIzY2OWZkMDsgc3Ryb2tlLXdpZHRoOiAyLjQ7IH0KICAgICAgLmV2aWRlbmNlIHsgZmlsbDogIzE3MjMxYzsgc3Ryb2tlOiAjNzJjNzdjOyBzdHJva2Utd2lkdGg6IDIuNDsgfQogICAgICAub3V0cHV0IHsgZmlsbDogIzFhMWQxYzsgc3Ryb2tlOiAjOGU5YTk0OyBzdHJva2Utd2lkdGg6IDI7IH0KICAgICAgLmJsdWUgeyBmaWxsOiBub25lOyBzdHJva2U6ICM3N2I3ZjI7IHN0cm9rZS13aWR0aDogMzsgc3Ryb2tlLWxpbmVjYXA6IHJvdW5kOyBzdHJva2UtbGluZWpvaW46IHJvdW5kOyB9CiAgICAgIC5ncmVlbiB7IGZpbGw6IG5vbmU7IHN0cm9rZTogIzc5Y2Y4Mzsgc3Ryb2tlLXdpZHRoOiAzOyBzdHJva2UtbGluZWNhcDogcm91bmQ7IHN0cm9rZS1saW5lam9pbjogcm91bmQ7IH0KICAgICAgLmdyYXkgeyBmaWxsOiBub25lOyBzdHJva2U6ICM2ODc1NmU7IHN0cm9rZS13aWR0aDogMi42OyBzdHJva2UtbGluZWNhcDogcm91bmQ7IHN0cm9rZS1saW5lam9pbjogcm91bmQ7IH0KICAgIDwvc3R5bGU+CiAgICA8bWFya2VyIGlkPSJhcnJvdy1ibHVlIiBtYXJrZXJXaWR0aD0iMTIiIG1hcmtlckhlaWdodD0iMTIiIHJlZlg9IjEwIiByZWZZPSI2IiBvcmllbnQ9ImF1dG8iIG1hcmtlclVuaXRzPSJ1c2VyU3BhY2VPblVzZSI+CiAgICAgIDxwYXRoIGQ9Ik0xIDFMMTAgNkwxIDExIiBmaWxsPSJub25lIiBzdHJva2U9IiM3N2I3ZjIiIHN0cm9rZS13aWR0aD0iMi4yIiBzdHJva2UtbGluZWNhcD0icm91bmQiIHN0cm9rZS1saW5lam9pbj0icm91bmQiLz4KICAgIDwvbWFya2VyPgogICAgPG1hcmtlciBpZD0iYXJyb3ctZ3JlZW4iIG1hcmtlcldpZHRoPSIxMiIgbWFya2VySGVpZ2h0PSIxMiIgcmVmWD0iMTAiIHJlZlk9IjYiIG9yaWVudD0iYXV0byIgbWFya2VyVW5pdHM9InVzZXJTcGFjZU9uVXNlIj4KICAgICAgPHBhdGggZD0iTTEgMUwxMCA2TDEgMTEiIGZpbGw9Im5vbmUiIHN0cm9rZT0iIzc5Y2Y4MyIgc3Ryb2tlLXdpZHRoPSIyLjIiIHN0cm9rZS1saW5lY2FwPSJyb3VuZCIgc3Ryb2tlLWxpbmVqb2luPSJyb3VuZCIvPgogICAgPC9tYXJrZXI+CiAgICA8bWFya2VyIGlkPSJhcnJvdy1ncmF5IiBtYXJrZXJXaWR0aD0iMTIiIG1hcmtlckhlaWdodD0iMTIiIHJlZlg9IjEwIiByZWZZPSI2IiBvcmllbnQ9ImF1dG8iIG1hcmtlclVuaXRzPSJ1c2VyU3BhY2VPblVzZSI+CiAgICAgIDxwYXRoIGQ9Ik0xIDFMMTAgNkwxIDExIiBmaWxsPSJub25lIiBzdHJva2U9IiM2ODc1NmUiIHN0cm9rZS13aWR0aD0iMiIgc3Ryb2tlLWxpbmVjYXA9InJvdW5kIiBzdHJva2UtbGluZWpvaW49InJvdW5kIi8+CiAgICA8L21hcmtlcj4KICA8L2RlZnM+CgogIDxyZWN0IHdpZHRoPSIxODAwIiBoZWlnaHQ9IjEwNDAiIGZpbGw9IiMxMDE0MTUiLz4KCiAgPHRleHQgeD0iODAiIHk9IjU4IiBjbGFzcz0ic2VjdGlvbiI+MSDCtyBMT0NBTCBFVklERU5DRSBTT1VSQ0VTPC90ZXh0PgogIDxyZWN0IHg9IjgwIiB5PSI4NiIgd2lkdGg9IjUwMCIgaGVpZ2h0PSIxMTIiIHJ4PSIxNiIgY2xhc3M9InNvdXJjZSIvPgogIDx0ZXh0IHg9IjMzMCIgeT0iMTMwIiBjbGFzcz0ibGFiZWwtc20iIHRleHQtYW5jaG9yPSJtaWRkbGUiPkZlYXR1cmUgVHJlZTwvdGV4dD4KICA8dGV4dCB4PSIzMzAiIHk9IjE2MyIgY2xhc3M9ImhpbnQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPkZlYXR1cmUgwrcgU3RvcnkgwrcgcmV2aWV3ZWQgcmVmZXJlbmNlczwvdGV4dD4KICA8cmVjdCB4PSI2NTAiIHk9Ijg2IiB3aWR0aD0iNTAwIiBoZWlnaHQ9IjExMiIgcng9IjE2IiBjbGFzcz0ic291cmNlIi8+CiAgPHRleHQgeD0iOTAwIiB5PSIxMzAiIGNsYXNzPSJsYWJlbC1zbSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+MTAgc2Vzc2lvbiBhZGFwdGVyczwvdGV4dD4KICA8dGV4dCB4PSI5MDAiIHk9IjE2MyIgY2xhc3M9ImhpbnQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnByb21wdHMgwrcgZGlhbG9ndWUgwrcgdG9vbCBjYWxsczwvdGV4dD4KICA8cmVjdCB4PSIxMjIwIiB5PSI4NiIgd2lkdGg9IjUwMCIgaGVpZ2h0PSIxMTIiIHJ4PSIxNiIgY2xhc3M9InNvdXJjZSIvPgogIDx0ZXh0IHg9IjE0NzAiIHk9IjEzMCIgY2xhc3M9ImxhYmVsLXNtIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5HaXQgKyBFbnRpcmU8L3RleHQ+CiAgPHRleHQgeD0iMTQ3MCIgeT0iMTYzIiBjbGFzcz0iaGludCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+Y29tbWl0cyDCtyBmaWxlcyDCtyBjaGVja3BvaW50czwvdGV4dD4KCiAgPHBhdGggZD0iTTMzMCAxOThWMjQ2SDkwMCIgY2xhc3M9ImdyYXkiLz4KICA8cGF0aCBkPSJNOTAwIDE5OFYyNDYiIGNsYXNzPSJncmF5Ii8+CiAgPHBhdGggZD0iTTE0NzAgMTk4VjI0Nkg5MDAiIGNsYXNzPSJncmF5Ii8+CiAgPHBhdGggZD0iTTkwMCAyNDZWMjc2IiBjbGFzcz0iYmx1ZSIgbWFya2VyLWVuZD0idXJsKCNhcnJvdy1ibHVlKSIvPgoKICA8dGV4dCB4PSI4MCIgeT0iMjc4IiBjbGFzcz0ic2VjdGlvbiI+MiDCtyBCT1VOREVEIExPQ0FMIFBST0pFQ1RJT048L3RleHQ+CiAgPHJlY3QgeD0iMTcwIiB5PSIzMDAiIHdpZHRoPSI0MzAiIGhlaWdodD0iMTE4IiByeD0iMTYiIGNsYXNzPSJwcm9jZXNzIi8+CiAgPHRleHQgeD0iMzg1IiB5PSIzNDUiIGNsYXNzPSJsYWJlbC1zbSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+Q29sbGVjdCBhIGJvdW5kZWQgd2luZG93PC90ZXh0PgogIDx0ZXh0IHg9IjM4NSIgeT0iMzc4IiBjbGFzcz0iaGludCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+dGltZSDCtyBzZXNzaW9ucyDCtyBjb21taXRzIMK3IHByb3ZpZGVyczwvdGV4dD4KICA8cGF0aCBkPSJNNjAwIDM1OUg3MDgiIGNsYXNzPSJibHVlIiBtYXJrZXItZW5kPSJ1cmwoI2Fycm93LWJsdWUpIi8+CiAgPHJlY3QgeD0iNzMwIiB5PSIzMDAiIHdpZHRoPSI0MzAiIGhlaWdodD0iMTE4IiByeD0iMTYiIGNsYXNzPSJwcm9jZXNzIi8+CiAgPHRleHQgeD0iOTQ1IiB5PSIzNDUiIGNsYXNzPSJsYWJlbC1zbSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+Tm9ybWFsaXplIHNhZmVseTwvdGV4dD4KICA8dGV4dCB4PSI5NDUiIHk9IjM3OCIgY2xhc3M9ImhpbnQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnJlZGFjdGlvbiDCtyByZWxhdGl2ZSBwYXRocyDCtyBhY3Rpb25zPC90ZXh0PgogIDxwYXRoIGQ9Ik0xMTYwIDM1OUgxMjY4IiBjbGFzcz0iYmx1ZSIgbWFya2VyLWVuZD0idXJsKCNhcnJvdy1ibHVlKSIvPgogIDxyZWN0IHg9IjEyOTAiIHk9IjMwMCIgd2lkdGg9IjM0MCIgaGVpZ2h0PSIxMTgiIHJ4PSIxNiIgY2xhc3M9InByb2Nlc3MiLz4KICA8dGV4dCB4PSIxNDYwIiB5PSIzNDUiIGNsYXNzPSJsYWJlbC1zbSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+UHJlc2VydmUgZ2FwczwvdGV4dD4KICA8dGV4dCB4PSIxNDYwIiB5PSIzNzgiIGNsYXNzPSJoaW50IiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5taXNzaW5nIHN0YXlzIHVub2JzZXJ2ZWQ8L3RleHQ+CgogIDxwYXRoIGQ9Ik05MDAgNDE4VjQ4NiIgY2xhc3M9ImdyZWVuIiBtYXJrZXItZW5kPSJ1cmwoI2Fycm93LWdyZWVuKSIvPgoKICA8dGV4dCB4PSI4MCIgeT0iNDc4IiBjbGFzcz0ic2VjdGlvbiI+MyDCtyBFVklERU5DRS1CT1VOREVEIENPUlJFTEFUSU9OPC90ZXh0PgogIDxyZWN0IHg9IjE4MCIgeT0iNTA2IiB3aWR0aD0iMTQ0MCIgaGVpZ2h0PSIxMzIiIHJ4PSIxOCIgY2xhc3M9ImV2aWRlbmNlIi8+CiAgPHRleHQgeD0iOTAwIiB5PSI1NTEiIGNsYXNzPSJsYWJlbCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+UmVsYXRlIGludGVudCwgc2Vzc2lvbnMsIHRvb2wgY2FsbHMsIHBhdGhzLCBhbmQgY29tbWl0czwvdGV4dD4KICA8dGV4dCB4PSIzNjAiIHk9IjU5OSIgY2xhc3M9ImxhYmVsLXNtIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5FeHBsaWNpdCAvIGRpcmVjdDwvdGV4dD4KICA8dGV4dCB4PSI3MjAiIHk9IjU5OSIgY2xhc3M9ImxhYmVsLXNtIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5PYnNlcnZlZCBzYW1lLXBhdGg8L3RleHQ+CiAgPHRleHQgeD0iMTA4MCIgeT0iNTk5IiBjbGFzcz0ibGFiZWwtc20iIHRleHQtYW5jaG9yPSJtaWRkbGUiPkNhbmRpZGF0ZTwvdGV4dD4KICA8dGV4dCB4PSIxNDQwIiB5PSI1OTkiIGNsYXNzPSJsYWJlbC1zbSIgdGV4dC1hbmNob3I9Im1pZGRsZSI+Q29udGV4dHVhbDwvdGV4dD4KCiAgPHBhdGggZD0iTTkwMCA2MzhWNzAwIiBjbGFzcz0iZ3JlZW4iIG1hcmtlci1lbmQ9InVybCgjYXJyb3ctZ3JlZW4pIi8+CgogIDx0ZXh0IHg9IjgwIiB5PSI2OTAiIGNsYXNzPSJzZWN0aW9uIj40IMK3IFJFQUQtT05MWSBJTlNQRUNUT1I8L3RleHQ+CiAgPHJlY3QgeD0iNTYwIiB5PSI3MjAiIHdpZHRoPSI2ODAiIGhlaWdodD0iODYiIHJ4PSIxOCIgY2xhc3M9ImV2aWRlbmNlIi8+CiAgPHRleHQgeD0iOTAwIiB5PSI3NzMiIGNsYXNzPSJsYWJlbCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+SGFybmVzc0luc3BlY3RvclJlcG9ydFYxPC90ZXh0PgogIDxwYXRoIGQ9Ik05MDAgODA2Vjg0NiIgY2xhc3M9ImdyYXkiIG1hcmtlci1lbmQ9InVybCgjYXJyb3ctZ3JheSkiLz4KCiAgPHJlY3QgeD0iNzAiIHk9Ijg3MCIgd2lkdGg9IjM4MCIgaGVpZ2h0PSIxMDAiIHJ4PSIxNSIgY2xhc3M9Im91dHB1dCIvPgogIDx0ZXh0IHg9IjI2MCIgeT0iOTEzIiBjbGFzcz0ibGFiZWwtc20iIHRleHQtYW5jaG9yPSJtaWRkbGUiPkRlbGl2ZXJ5IFRyZWUgLyBEYXRlPC90ZXh0PgogIDx0ZXh0IHg9IjI2MCIgeT0iOTQ1IiBjbGFzcz0iaGludCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+Y2hvb3NlIHRoZSByZXZpZXcgc2NvcGU8L3RleHQ+CiAgPHJlY3QgeD0iNDkwIiB5PSI4NzAiIHdpZHRoPSIzODAiIGhlaWdodD0iMTAwIiByeD0iMTUiIGNsYXNzPSJvdXRwdXQiLz4KICA8dGV4dCB4PSI2ODAiIHk9IjkxMyIgY2xhc3M9ImxhYmVsLXNtIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5UaHJlZSBldmlkZW5jZSBsYW5lczwvdGV4dD4KICA8dGV4dCB4PSI2ODAiIHk9Ijk0NSIgY2xhc3M9ImhpbnQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPnByb21wdCDCtyBhY3Rpdml0eSDCtyBkZWxpdmVyeTwvdGV4dD4KICA8cmVjdCB4PSI5MTAiIHk9Ijg3MCIgd2lkdGg9IjM4MCIgaGVpZ2h0PSIxMDAiIHJ4PSIxNSIgY2xhc3M9Im91dHB1dCIvPgogIDx0ZXh0IHg9IjExMDAiIHk9IjkxMyIgY2xhc3M9ImxhYmVsLXNtIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5FdmlkZW5jZSBEcmF3ZXI8L3RleHQ+CiAgPHRleHQgeD0iMTEwMCIgeT0iOTQ1IiBjbGFzcz0iaGludCIgdGV4dC1hbmNob3I9Im1pZGRsZSI+ZmFjdHMgwrcgc291cmNlIMK3IGxpbWl0YXRpb25zPC90ZXh0PgogIDxyZWN0IHg9IjEzMzAiIHk9Ijg3MCIgd2lkdGg9IjQwMCIgaGVpZ2h0PSIxMDAiIHJ4PSIxNSIgY2xhc3M9Im91dHB1dCIvPgogIDx0ZXh0IHg9IjE1MzAiIHk9IjkxMyIgY2xhc3M9ImxhYmVsLXNtIiB0ZXh0LWFuY2hvcj0ibWlkZGxlIj5TZXNzaW9uIFRyYWNlIC8gUmVwbGF5PC90ZXh0PgogIDx0ZXh0IHg9IjE1MzAiIHk9Ijk0NSIgY2xhc3M9ImhpbnQiIHRleHQtYW5jaG9yPSJtaWRkbGUiPm9ic2VydmU7IG5ldmVyIHJlcnVuIG9yIHJlc3VtZTwvdGV4dD4KCiAgPHBhdGggZD0iTTkwMCA4NDZIMjYwVjg3MCIgY2xhc3M9ImdyYXkiLz4KICA8cGF0aCBkPSJNOTAwIDg0Nkg2ODBWODcwIiBjbGFzcz0iZ3JheSIvPgogIDxwYXRoIGQ9Ik05MDAgODQ2SDExMDBWODcwIiBjbGFzcz0iZ3JheSIvPgogIDxwYXRoIGQ9Ik05MDAgODQ2SDE1MzBWODcwIiBjbGFzcz0iZ3JheSIvPgo8L3N2Zz4K" width="1800" height="1040" class="img_ev3q"></p>
<ul>
<li class=""><strong>Intent</strong> is the semantic starting point of a change — a requirement, an Issue, a Spec, or an architectural constraint.</li>
<li class=""><strong>Process</strong> is how the change actually unfolds. For an agent, that is mostly the session record and the searching, reading, editing, and verifying inside it.</li>
<li class=""><strong>Output</strong> is the final result the agent delivers into the engineering system. For now, the clearest anchor is the code commit.</li>
</ul>
<p>So Story, Session, and Commit are not three parallel abstractions. They are the <strong>observable objects of Intent, Process, and Output in today's software toolchain</strong> (intent may also appear as an Issue or a Spec depending on the context). What Harness Inspector does is not to render a conversation more completely, but to rebuild a <em>traceable</em> delivery chain from requirement to commit.</p>
<p>As a narrative it reads like a straight line. In a real project it looks more like an evidence graph: one Story may span several sessions, and one session may touch several commits.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-small-example-from-story-to-commit">A small example: from Story to Commit<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#a-small-example-from-story-to-commit" class="hash-link" aria-label="Direct link to A small example: from Story to Commit" title="Direct link to A small example: from Story to Commit" translate="no">​</a></h3>
<p>We published a read-only public sample on the Better Harness docs site (English demo data, no local content is read): <a href="https://qoderai.github.io/better-harness/inspector" target="_blank" rel="noopener noreferrer" class="">https://qoderai.github.io/better-harness/inspector</a>.</p>
<p>Looking at the session alone, you see a stream of search, read, edit, and test activity. It is hard to be sure it stayed anchored to the original requirement, and impossible to know which of those edits reached the repository. Looking at the commit alone, you see which files changed, but not how the agent understood the problem, built context, or verified the result before committing.</p>
<p>Put Story, Session, and Commit in one view, and the change finally becomes a reasonably complete delivery: the Story says <em>why</em> the change was needed, the Session shows <em>how</em> it happened, and the Commit records <em>what</em> was left behind.</p>
<p>Connecting requirement, agent behavior, and commit let us inspect a delivery as a whole for the first time. But once we opened a few real sessions with hundreds of tool calls, another problem showed up fast: <strong>being able to connect a delivery is not the same as being able to read it.</strong></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-harness-inspector-reads-an-agent-delivery">How Harness Inspector reads an agent delivery<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#how-harness-inspector-reads-an-agent-delivery" class="hash-link" aria-label="Direct link to How Harness Inspector reads an agent delivery" title="Direct link to How Harness Inspector reads an agent delivery" translate="no">​</a></h2>
<p>Around the Better Harness harness model, we define Harness Inspector like this:</p>
<blockquote>
<p><strong>Harness Inspector is a local, read-only workbench for agent deliveries. It brings requirements, agent sessions, file activity, and Git commits into one interactive view, so you can inspect why a software change happened, how it happened, and what it left behind.</strong></p>
</blockquote>
<p>Centered on a single complete delivery, it offers three ways to observe the chain from requirement to commit:</p>
<ul>
<li class=""><strong>Workbench</strong> — the relationships between requirement, session, and commit.</li>
<li class=""><strong>Trace</strong> — the internal structure of a session.</li>
<li class=""><strong>Replay</strong> — the task replayed in event order.</li>
</ul>
<p>Put simply: <strong>Workbench for relationships, Trace for structure, Replay for order.</strong> Together they reconstruct how a requirement passes through an agent's understanding, exploration, editing, and verification into a reviewable commit.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="workbench-connecting-intent-process-and-output">Workbench: connecting intent, process, and output<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#workbench-connecting-intent-process-and-output" class="hash-link" aria-label="Direct link to Workbench: connecting intent, process, and output" title="Direct link to Workbench: connecting intent, process, and output" translate="no">​</a></h3>
<p>Workbench is the whole-delivery view. On the left is the user requirement that triggered the session, plus the goal refinements made along the way. In the middle is what happened during the session — searches, reads, tool calls, and Git operations. On the right are the commits observed within the current scope and the files they changed.</p>
<p>Its point is not to pile data together, but to show the relationships between requirement, process, and output that have <strong>actually been observed</strong>. When the evidence for a relationship is weak, Inspector keeps it as a <em>candidate</em> or leaves it <em>unmapped</em>, rather than auto-assembling a delivery path that only looks complete.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="trace-reading-a-session-as-a-work-trajectory">Trace: reading a session as a work trajectory<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#trace-reading-a-session-as-a-work-trajectory" class="hash-link" aria-label="Direct link to Trace: reading a session as a work trajectory" title="Direct link to Trace: reading a session as a work trajectory" translate="no">​</a></h3>
<p>Workbench helps you find a delivery; Trace expands the session inside it.</p>
<p>Inside a session, Trace organizes user input, intermediate responses, tool calls, and file activity by Turn, and connects their positions in time along the timeline at the top. Click a segment to jump to the corresponding call; consecutive repeated activity is folded so a flood of similar operations does not drown out the changes that matter.</p>
<p>Trace does not try to recover reasoning the model never exposed. It reorganizes the behavior that <em>was</em> recorded into a readable work trajectory, so you can inspect how the agent searched for context, edited code, and ran verification.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="replay-revisiting-a-delivery-in-event-order">Replay: revisiting a delivery in event order<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#replay-revisiting-a-delivery-in-event-order" class="hash-link" aria-label="Direct link to Replay: revisiting a delivery in event order" title="Direct link to Replay: revisiting a delivery in event order" translate="no">​</a></h3>
<p>Replay walks through the retained events step by step. A reviewer can move through user input, agent response, tool call, file, and commit in order, watching where the agent formed a direction and when it made and verified a change.</p>
<p>It is a read-only replay of evidence only. It does not rerun tools, restore the workspace, or resume the original session; where an exact timestamp was never recorded, it keeps order only and invents nothing that was not observed.</p>
<p>Workbench establishes the delivery context, Trace expands the session's work trajectory, and Replay restores the order of events. Together they turn a requirement-to-commit agent delivery into something you can enter and inspect layer by layer. And once a delivery reads clearly, we can return to the original question: which of the paths in this trajectory are actually worth keeping.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="from-delivery-evidence-to-skill-distillation">From delivery evidence to SKILL distillation<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#from-delivery-evidence-to-skill-distillation" class="hash-link" aria-label="Direct link to From delivery evidence to SKILL distillation" title="Direct link to From delivery evidence to SKILL distillation" translate="no">​</a></h2>
<p>Once a delivery can be seen clearly, we can go back to where we started: which experiences inside a session are worth distilling into a SKILL? Clearly not the most frequent tool call. An agent may reread the same file over and over because its context was thin, or retry a failing command repeatedly — behavior that is frequent but is more likely noise than reusable engineering experience.</p>
<p>What is worth keeping is usually a work path that recurs across similar tasks <em>and</em> is supported by the final output and verification result: how the agent scoped the change from the requirement, how it built the necessary context, and how it completed the edit, ran verification, and checked the result. Only by placing those actions back into their Story, Session, and Commit can we tell whether they were an incidental choice for this one task or a relatively stable way of working that transfers to others.</p>
<p>So automatic SKILL distillation is not summarizing one session into a new <code>SKILL.md</code>. It is recognizing stable patterns across many real deliveries, then giving each one its applicable scenario, context boundaries, execution steps, and verification method. What Inspector solves today is the most basic link in that chain: letting real deliveries leave behind bounded, inspectable evidence. Only on that foundation can we compare similar tasks, form SKILL candidates, and verify in later deliveries whether they actually improved how the agent works.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="closing-see-the-delivery-first-then-distill-the-experience">Closing: see the delivery first, then distill the experience<a href="https://qoderai.github.io/better-harness/blog/harness-inspector#closing-see-the-delivery-first-then-distill-the-experience" class="hash-link" aria-label="Direct link to Closing: see the delivery first, then distill the experience" title="Direct link to Closing: see the delivery first, then distill the experience" translate="no">​</a></h2>
<p>What we set out to solve was how to recognize reusable work paths inside agent sessions. But once you actually dig in, a session can only record what the agent did — it cannot, on its own, explain why the task began, or prove which actions became an engineering artifact.</p>
<p>Harness Inspector reconnects requirement, session, file activity, and Git commit, making a single agent delivery observable, inspectable, and traceable. Only by seeing a delivery clearly can we judge which behavior was merely incidental and which path is worth distilling into a SKILL.</p>
<p>From session to delivery, and from delivery to distilled experience — that is the starting point of automatic SKILL evolution.</p>]]></content:encoded>
            <category>better-harness</category>
            <category>harness-inspector</category>
            <category>coding-agent</category>
            <category>provenance</category>
            <category>agent-skills</category>
        </item>
        <item>
            <title><![CDATA[How Agent Plugins Become Engineered: Five Practices from Better Harness]]></title>
            <link>https://qoderai.github.io/better-harness/blog/agent-plugin-engineering</link>
            <guid>https://qoderai.github.io/better-harness/blog/agent-plugin-engineering</guid>
            <pubDate>Sun, 09 Aug 2026 06:10:00 GMT</pubDate>
            <description><![CDATA[Writing a capability into SKILL.md is easy. Keeping it discoverable, executable, and provably better across users, projects, and hosts is an engineering problem.]]></description>
            <content:encoded><![CDATA[<p>Writing a capability into a <code>SKILL.md</code> and packaging it as an Agent plugin is
not hard. The hard part comes later: once that capability is invoked over and
over by different users, in different projects, and on different Agent hosts,
how do you guarantee that it is still triggered, executed, and verified
correctly? And how do you prove that a change made it better rather than worse?</p>
<p>Drawing on how Better Harness is actually developed, this post walks through
five engineering practices - spec-driven behavior, context orchestration,
deterministic verification, behavioral evaluation, and the evidence loop - that
move a plugin capability from "it works when I run it" to a software asset that
is verifiable, maintainable, and safe to evolve.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="prologue-starting-from-agent-plugins">Prologue: starting from Agent plugins<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#prologue-starting-from-agent-plugins" class="hash-link" aria-label="Direct link to Prologue: starting from Agent plugins" title="Direct link to Prologue: starting from Agent plugins" translate="no">​</a></h2>
<p>Over the past three months, alongside designing the Canvas-based Agentic UI
system, our other ongoing investment at Qoder has been building the Agent plugin
ecosystem. From curating existing software-engineering workflow plugins in the
community to creating plugins such as Architecture Visualization and Design
Review, we have been trying to package capabilities that used to live scattered
across different roles and different tools into something an Agent can discover
and call.</p>
<p>In the Architecture Visualization plugin, for example, the Agent can render the
dependency relationships between architecture modules directly on Canvas and
share the result with teammates so they can understand the structure of the
system.</p>
<p>At the beginning, a plugin felt mostly like an extension of the Agent's
capability boundary: whatever was missing, we added a Skill, a tool, or a
workflow for it. But as the plugins grew in number and complexity, another
question surfaced:</p>
<p>How should these capabilities be maintained over the long run?</p>
<p>Building Better Harness made that question impossible to avoid. Better Harness
targets workflow analysis and continuous improvement for coding agents, and it
contains a core Skill, a large amount of JavaScript, reference material,
evaluation logic, and adapters for different Agent hosts.</p>
<p>Once those capabilities became part of a plugin that has to be maintained over
time, we realized the problem was no longer "how do I write a good <code>SKILL.md</code>".</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="what-exactly-is-an-agent-plugin">What exactly is an Agent Plugin?<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#what-exactly-is-an-agent-plugin" class="hash-link" aria-label="Direct link to What exactly is an Agent Plugin?" title="Direct link to What exactly is an Agent Plugin?" translate="no">​</a></h3>
<blockquote>
<p>If you are already familiar with Agent Plugins and Agent Skills, skip to the
next section.</p>
</blockquote>
<p>Put simply, an Agent Plugin is the delivery and distribution boundary of a
capability, and a Skill is a unit inside it that an Agent can discover and
execute independently. A typical plugin can contain Skills, MCP servers, hooks,
and host-specific extensions at the same time:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">my-plugin/</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">├── plugin.json</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">├── skills/</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">│   └── summarize/</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">│       ├── SKILL.md</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">│       ├── scripts/</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">│       └── references/</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">├── mcp.json</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">└── com.example.client/</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">    └── hooks/</span><br></div></code></pre></div></div>
<p>A <code>SKILL.md</code> is not the plugin itself; it is closer to an entry point for one
capability inside it. It describes when the capability should be discovered, how
it should be executed, and which additional material and tools need to be read
next. If a Skill is just a prompt you occasionally use yourself, that
distinction hardly matters. The moment it enters a plugin and gets invoked
repeatedly by different users, projects, and hosts, it does.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-do-plugins-need-engineering">Why do plugins need engineering?<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#why-do-plugins-need-engineering" class="hash-link" aria-label="Direct link to Why do plugins need engineering?" title="Direct link to Why do plugins need engineering?" translate="no">​</a></h3>
<p>For personal use, writing down the task and the steps is usually enough. Once a
plugin is called repeatedly across users, projects, and hosts, a set of problems
that look a lot like software engineering appear on their own:</p>
<ul>
<li class=""><strong>Behavioral contract: how is it supposed to work?</strong> Which situations should
trigger it and which should not; what context does it need, which tools may it
call, and which behaviors are forbidden?</li>
<li class=""><strong>Change verification: is it still correct after a change?</strong> When a Skill, a
script, or a reference document changes, how do you confirm that the existing
capability was not broken and that a fixed problem does not come back?</li>
<li class=""><strong>Environment compatibility: does it still work elsewhere?</strong> When the model,
the Agent, the host, or the runtime changes, can the capability still be
discovered, loaded, and executed correctly?</li>
<li class=""><strong>Outcome assessment: did it actually make the Agent better?</strong> Even if the
Skill is triggered correctly and executed fully, how do you show that it
improved the result of a real task compared with not using it?</li>
</ul>
<p>None of these are solved by "writing a more detailed prompt". They map onto very
familiar software-engineering problems: defining contracts, verifying changes,
managing compatibility, and assessing real effect. Take them further and you
land on specifications, knowledge and dependency organization, interface and
permission boundaries, automated tests, regression verification, and
cross-model, cross-host behavioral evaluation.</p>
<p>It was during the development of Better Harness that we gradually started to
understand Skills differently:</p>
<blockquote>
<p>Once a Skill leaves the personal-prompt stage and enters a plugin ecosystem,
the problems it faces look increasingly like software, not prompting.</p>
</blockquote>
<p>The five engineering practices below follow from that shift.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-spec-driven-write-agent-behavior-as-a-verifiable-contract">1. Spec-driven: write Agent behavior as a verifiable contract<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#1-spec-driven-write-agent-behavior-as-a-verifiable-contract" class="hash-link" aria-label="Direct link to 1. Spec-driven: write Agent behavior as a verifiable contract" title="Direct link to 1. Spec-driven: write Agent behavior as a verifiable contract" translate="no">​</a></h2>
<p>Spec-driven development is the part we started practicing earliest in Better
Harness. The core idea is to write down "when to act, what to do, and how well
it must be done" as a verifiable contract before implementation.</p>
<p>Our <a href="https://github.com/QoderAI/better-harness/blob/main/AGENTS.md" target="_blank" rel="noopener noreferrer" class=""><code>AGENTS.md</code></a>
states it directly:</p>
<blockquote>
<p>For non-trivial behavior changes to Skills, scripts, templates, host adapters,
or review workflows, establish a spec and traceable acceptance criteria first.</p>
</blockquote>
<p>A spec has to answer at least a few questions: which requests should trigger the
capability and which similar-sounding requests should not; what information must
be read and which tools may be called; which behaviors are explicitly
forbidden; whether the Agent should stop, degrade, or hand back to a human when
evidence is insufficient; and finally, what evidence proves the implementation
matches the intent.</p>
<p>Every acceptance criterion should carry a stable id such as AC-01 or AC-02, each
mapped to implementation, test, and review evidence. One of our boundary-hardening
specs, for instance, decomposes reproduced problems into AC-01 through AC-09.</p>
<p>Concrete examples: when the Agent scans written content, it must not echo
secrets; when analyzing one workspace, it must not quietly count sessions from
other directories; when the Git baseline cannot be established, it must stop
explicitly instead of disguising failure as "no changes". AC-09 then closes the
loop with a full test and packaging verification.</p>
<p>Specs themselves need review. In the
<a href="https://github.com/QoderAI/better-harness/blob/main/.agents/skills/triangulate-spec-review/SKILL.md" target="_blank" rel="noopener noreferrer" class=""><code>triangulate-spec-review</code></a>
Skill, we have at least two - usually three - review Agents inspect the same
context from different angles: implementation complexity, ease of use, and
long-term evolution. A lead Agent merges duplicate findings, checks evidence, and
edits the document; the other reviewers only provide independent judgment and do
not touch files. The value of being spec-driven is exactly this: ambiguity moves
from run time to design time.</p>
<p>Spec-driven work has a boundary too. When acceptance criteria are written too
finely, AI tends to turn tests into word-by-word matching against the Skill text
instead of verifying real behavior, which makes the tests less stable. A spec
should constrain observable behavior, not freeze specific wording.</p>
<p>A spec defines how a capability should work. The next step is making sure the
Agent can actually find the knowledge that supports that behavior while it runs.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-context-orchestration-make-skill-knowledge-arrive-when-it-is-needed">2. Context orchestration: make Skill knowledge arrive when it is needed<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#2-context-orchestration-make-skill-knowledge-arrive-when-it-is-needed" class="hash-link" aria-label="Direct link to 2. Context orchestration: make Skill knowledge arrive when it is needed" title="Direct link to 2. Context orchestration: make Skill knowledge arrive when it is needed" translate="no">​</a></h2>
<p>Agent Skills generally rely on progressive disclosure: knowledge should not all
enter the context at once, but unfold as the task requires.</p>
<blockquote>
<p>The host first reads <code>name</code> and <code>description</code> to complete discovery, loads the
full <code>SKILL.md</code> only after deciding to use the Skill, and pulls finer material
into context as the task demands. This is essentially context engineering for
Agents: rather than pushing all knowledge into the model at once, you design
when knowledge appears, where it enters from, and how it stays traceable.</p>
</blockquote>
<p>Translated into directory structure, the entry point should stay short:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">my-skill/</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">├── SKILL.md          # trigger conditions, main flow, stop conditions</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">├── references/       # judgment rules loaded on demand</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">├── scripts/          # deterministic logic that can be re-run</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">└── assets/           # templates and delivery skeletons</span><br></div></code></pre></div></div>
<p><code>SKILL.md</code> owns triggering, routing, and stop conditions; detailed judgment goes
into <code>references/</code>; stable executable logic goes into <code>scripts/</code>; templates and
delivery skeletons go into <code>assets/</code>.</p>
<p>In practice, though, we found that splitting knowledge apart is not enough: a
file existing does not mean the Agent can find it, and having been referenced
once does not mean the path still resolves after a refactor.</p>
<p>So we protect knowledge routing along three axes:</p>
<ul>
<li class=""><strong>Discoverable:</strong> bring relevant material into the Agent's knowledge path
through explicit entry points and references in <code>SKILL.md</code>, rather than relying
on ad-hoc search.</li>
<li class=""><strong>Reachable:</strong> use
<a href="https://github.com/QoderAI/better-harness/blob/main/test/skills-docs/doc-link-graph.test.mjs" target="_blank" rel="noopener noreferrer" class="">automated tests</a>
to check that relative links resolve and that every document a Skill needs is
genuinely routed from its entry point.</li>
<li class=""><strong>Traceable:</strong> generate a Mermaid graph from the real Markdown references with
a <a href="https://github.com/QoderAI/better-harness/blob/main/scripts/doc-link-graph/cli.mjs" target="_blank" rel="noopener noreferrer" class="">doc-link-graph generator</a>,
and verify that the generated output still matches the current reference
relationships.</li>
</ul>
<p>That way the Markdown references themselves are the source of truth, and the
graph is only a verifiable projection of knowledge routing. When documents move,
links break, or routing changes, a machine notices in time instead of relying on
maintainers to sync manually.</p>
<p>The goal of context orchestration is not to make the Agent read more, but to
make the right knowledge enter the context at the right moment through a path
that still works. Once knowledge arrives reliably, the next question is which
boundaries should be decided by a program rather than guessed by a model.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-deterministic-verification-give-the-machine-what-it-can-decide">3. Deterministic verification: give the machine what it can decide<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#3-deterministic-verification-give-the-machine-what-it-can-decide" class="hash-link" aria-label="Direct link to 3. Deterministic verification: give the machine what it can decide" title="Direct link to 3. Deterministic verification: give the machine what it can decide" translate="no">​</a></h2>
<p>Not every question inside a Skill or a plugin should be handed to the model.
Keywords, regular expressions, and program checks are well suited to problems
with clear boundaries and decidable outcomes - metadata, directory structure,
broken links, data formats, deprecated names, permission declarations. What they
cannot do is prove that the Agent understood the task, and they are no substitute
for semantic quality review.</p>
<p>In Better Harness we hold those boundaries with three deterministic layers:</p>
<ul>
<li class=""><strong>Lint:</strong> check file headers, directory structure, broken links, data formats,
and permission declarations, catching cheap and decidable problems early.</li>
<li class=""><strong>Unit tests:</strong> verify deterministic logic such as scripts, parsers, and
template transforms, so basic capabilities still hold after a change.</li>
<li class=""><strong>Contract tests:</strong> verify what a command or entry point is actually allowed to
do, including inputs and outputs, error states, artifact locations, and
side-effect boundaries.</li>
</ul>
<p>A representative example is how Better Harness tests the <code>--help</code> path.
<a href="https://github.com/QoderAI/better-harness/blob/main/test/cli/better-harness-cli.test.mjs" target="_blank" rel="noopener noreferrer" class=""><code>better-harness-cli.test.mjs</code></a>
does not merely check that the help text is correct; it further verifies that
running a help command must not read the workspace, write files, wait on standard
input, spawn a child process, or access the network. If any one of those side
effects occurs, the test fails.</p>
<p>What is really being verified here is not what the help text looks like, but what
this entry point is and is not permitted to do. That is the value of
deterministic verification: boundaries that model behavior can easily paper over
become engineering constraints that are checked automatically and regressed
continuously. In an Agent system, models are better at understanding, planning,
and trading off; deterministic programs are better at verifying, constraining,
and refusing.</p>
<blockquote>
<p>If a program can decide it, do not make the model guess. If it needs semantic
judgment, do not force it into a string assertion.</p>
</blockquote>
<p>Deterministic verification only guards decidable boundaries, though. It cannot
prove the Agent actually follows the Skill in a real task. That requires
behavioral evaluation.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-behavioral-evaluation-verify-the-agent-really-does-it">4. Behavioral evaluation: verify the Agent really does it<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#4-behavioral-evaluation-verify-the-agent-really-does-it" class="hash-link" aria-label="Direct link to 4. Behavioral evaluation: verify the Agent really does it" title="Direct link to 4. Behavioral evaluation: verify the Agent really does it" translate="no">​</a></h2>
<p>A Skill existing, being discovered, being loaded, being executed, and finally
producing an improvement are five different things. Passing static checks only
says that the files and scripts have no obvious defects. Whether the Agent
chooses the Skill at the right moment, performs the key steps, and stops when
evidence is insufficient still needs behavioral evaluation.</p>
<p>In Better Harness we usually prepare three kinds of scenarios: positive cases
that should trigger, negative cases that should not, and boundary cases worded
similarly but with a different intent. During execution we watch four things:</p>
<ul>
<li class=""><strong>Selection:</strong> was it used when it should be, and not misfired when it should
not be?</li>
<li class=""><strong>Context:</strong> did it read the material it genuinely needed?</li>
<li class=""><strong>Execution:</strong> did the key steps, tools, and permissions stay within the
constraints?</li>
<li class=""><strong>Outcome:</strong> can the final artifact pass independent verification?</li>
</ul>
<p>The same scenario also needs repeated runs, because one success only proves that
this particular run worked, not that the behavior is stable.</p>
<p>At the host boundary, Better Harness additionally runs end-to-end verification
through the real Qoder CLI and plugin loading chain, launched from a neutral
directory:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">python3 &lt;skill-creator-root&gt;/scripts/quick_validate.py &lt;skill-dir&gt;</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">qodercli --cwd &lt;neutral-dir&gt; --plugin-dir &lt;plugin-root&gt; -p "&lt;forward-test-prompt&gt;"</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">qodercli plugin validate &lt;plugin-root&gt;</span><br></div></code></pre></div></div>
<p>Starting from a neutral directory prevents the Agent from accidentally borrowing
configuration, context, or undeclared dependencies from the current repository,
which would make "it works" look more optimistic than reality. At the same time,
whatever <code>qodercli</code> answers counts only as model-behavior evidence; whether the
run passes is still decided by deterministic evidence such as the local
validator, artifact inspection, and Git status.</p>
<p>Simulated cases cover exceptional and boundary situations; real host tests verify
whether the chain from plugin loading to Skill discovery, material reading, tool
invocation, and final delivery is genuinely connected. What behavioral evaluation
sets out to prove is not that the Agent has <em>seen</em> the Skill, but that the key
behaviors the Skill requires actually happened.</p>
<p>Proving that the Agent executed the Skill, however, is still not proof that the
Skill produced a better result. That belongs to the evidence loop.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="5-evidence-loop-prove-the-skill-works-before-talking-about-self-evolution">5. Evidence loop: prove the Skill works before talking about self-evolution<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#5-evidence-loop-prove-the-skill-works-before-talking-about-self-evolution" class="hash-link" aria-label="Direct link to 5. Evidence loop: prove the Skill works before talking about self-evolution" title="Direct link to 5. Evidence loop: prove the Skill works before talking about self-evolution" translate="no">​</a></h2>
<p>Behavioral evaluation answers "did the Agent follow the Skill". The evidence loop
goes one step further: given the same task, does the Agent actually do better
<em>with</em> this Skill?</p>
<p>In Better Harness this is designed as an
<a href="https://github.com/QoderAI/better-harness/blob/main/references/agent-customize/skill-eval.md" target="_blank" rel="noopener noreferrer" class="">evaluation execution protocol</a>.
Holding the task, model, tools, permissions, and test environment constant, we
run three arms: no Skill, current Skill version, candidate Skill version.</p>
<p>An evaluation looks at more than the final success rate. It also observes whether
the key steps were executed, whether verification was complete, the time and
token cost, and whether extra side effects appeared. Otherwise a Skill that looks
"better" may simply be using more permissions or more resources.</p>
<p>One failure mode deserves special attention: fake usage, where the Agent finds
the Skill and reads <code>SKILL.md</code> but never performs the steps it requires. Better
Harness calls this <strong>routed-but-not-applied</strong>.</p>
<p>So we keep judging along an evidence chain:</p>
<ul>
<li class=""><strong>Does it exist?</strong> Are the Skill and its supporting mechanisms present?</li>
<li class=""><strong>Can it be found?</strong> Can a real task discover and select it?</li>
<li class=""><strong>Was it executed?</strong> Did the required key steps actually happen?</li>
<li class=""><strong>Is it effective?</strong> Compared with not using the Skill, did the result improve?</li>
</ul>
<p>In the <a href="https://github.com/QoderAI/better-harness/blob/main/models/agent-work-loop.md" target="_blank" rel="noopener noreferrer" class="">Agent Work Loop</a>
this maps to <strong>Present → Wired → Exercised → Outcome-supported</strong>. There is a
single governing principle: a conclusion may only go as far as the evidence goes.</p>
<p>That also matches where recent Skill-evaluation research is heading. SkillsBench
focuses on the outcome difference between using and not using a Skill on the same
task; Skill Coverage further checks whether the behaviors a Skill requires really
appear in the execution trace. The former answers "did the result get better",
the latter "did the process actually happen".</p>
<p>Outcome improvement and process coverage are both required; neither alone is
enough. Only after that evidence chain is in place does self-evolution become
meaningful. Otherwise Trace2Skill, EvoSkill, CoEvoSkills, or SkillOpt may just
help an Agent produce more unverified Skills faster.</p>
<p>Skill evolution should not start from "generate more experience". It should start
from proving that this change really made the next run better.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="conclusion">Conclusion<a href="https://qoderai.github.io/better-harness/blog/agent-plugin-engineering#conclusion" class="hash-link" aria-label="Direct link to Conclusion" title="Direct link to Conclusion" translate="no">​</a></h2>
<p>The core of Agent plugin engineering is not writing more elaborate Skills. It is
treating the capabilities inside a plugin as software assets: constrain behavior
with specs, organize knowledge with context orchestration, hold boundaries with
deterministic verification, confirm execution with behavioral evaluation, and
judge effect with controlled comparison.</p>
<p>When those mechanisms work together, a plugin finally moves from "it works when I
run it" to "verifiable, maintainable, and safe to evolve".</p>
<p>Developing Agent plugins as software also means that every change should leave
behind enough evidence to show that it got better.</p>]]></content:encoded>
            <category>better-harness</category>
            <category>agent-plugin</category>
            <category>agent-skills</category>
            <category>spec-driven</category>
            <category>evaluation</category>
        </item>
        <item>
            <title><![CDATA[/better-harness Goes Open Source]]></title>
            <link>https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source</link>
            <guid>https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source</guid>
            <pubDate>Thu, 30 Jul 2026 10:00:00 GMT</pubDate>
            <description><![CDATA[We are open-sourcing the engineering practices, evidence model, and runnable workflow behind Better Harness for coding agents.]]></description>
            <content:encoded><![CDATA[<p>Last week, we built Better Harness into Qoder Desktop. After launch, many users
asked the same question: <strong>Will this be open source?</strong></p>
<p>In its first three days, 100,000 people tried Better Harness.</p>
<p>The answer is yes.</p>
<p>Today, Better Harness is officially open source. You can find the project at
<a href="https://github.com/QoderAI/better-harness" target="_blank" rel="noopener noreferrer" class="">github.com/QoderAI/better-harness</a>.</p>
<p>Better Harness is an open-source analysis and continuous-improvement tool for
coding-agent workflows. It connects the engineering practices, evaluation
model, and runtime capabilities of Harness Engineering and Loop Engineering.
The initial open-source release supported Claude Code, Codex, Qoder, and Cursor
with one shared judgment model, although session analysis, evidence coverage,
and output capabilities were not yet identical across the four hosts. Qoder,
which had already been exercised repeatedly in real development workflows, was
the most complete reference implementation at launch.</p>
<table><thead><tr><th>Launch host</th><th>Installation or loading path</th><th>Default output</th></tr></thead><tbody><tr><td><strong>Claude Code</strong></td><td>Add the repository marketplace, then install the plugin</td><td>HTML + Markdown</td></tr><tr><td><strong>Codex</strong></td><td>Install through a Git marketplace</td><td>HTML + Markdown</td></tr><tr><td><strong>Qoder Desktop</strong></td><td>Built in; no separate installation</td><td>Canvas</td></tr><tr><td><strong>Cursor Agent</strong></td><td>Load from source</td><td>HTML / Markdown</td></tr></tbody></table>
<blockquote>
<p><strong>Editor's note:</strong> This table records the launch state. Better Harness now
publishes additional host integrations, and entrypoints differ by host. See
the current <a href="https://qoderai.github.io/better-harness/docs/installation" target="_blank" rel="noopener noreferrer" class="">Installation guide</a>
before installing or running a review.</p>
</blockquote>
<p>At launch, the shared workflow was commonly invoked as:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">/better-harness</span><br></div></code></pre></div></div>
<p>You could start the analysis, inspect the report as it ran, and continue
reading this article. When a category of evidence was unavailable, Better
Harness preserved that boundary in the result instead of substituting config
counts or data from another host.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="better-harness-cares-about-what-the-agent-did-in-the-task">Better Harness cares about what the agent did in the task<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#better-harness-cares-about-what-the-agent-did-in-the-task" class="hash-link" aria-label="Direct link to Better Harness cares about what the agent did in the task" title="Direct link to Better Harness cares about what the agent did in the task" translate="no">​</a></h2>
<p>Imagine that an agent modifies a module, runs one test command, and declares
the task complete. The real question is not whether the repository <em>has</em> tests.
It is whether that test was relevant to the change, covered the main risks, and
produced enough evidence to support delivery.</p>
<p>Better Harness therefore does not treat the presence of AGENTS.md, Rules,
Skills, MCP servers, Hooks, memories, tests, or CI as proof that they influenced
the task. Every improvement item—a Finding—must include traceable evidence, a
specific user impact, the smallest repair boundary, and a way to verify the
result after the repair.</p>
<p>In one self-analysis snapshot produced from Codex, Better Harness did not turn
the absence of an executed Codex host test into the stronger claim that
“Codex has failed.” It kept the unexecuted host test as an explicit evidence
boundary. Scores can help locate a problem, but the conclusion, impact, repair
scope, and verification method are what matter.</p>
<p>That is the difference between Better Harness and a configuration checklist. A
checklist tells you what the project possesses. Better Harness asks whether
those capabilities actually helped the agent complete a trustworthy task.</p>
<p>And if the judgments in a report are meant to be inspected, changed, and
reverified, open source cannot stop at publishing an executable entrypoint.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-three-layer-open-source-system-behind-better-harness">The three-layer open-source system behind Better Harness<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#the-three-layer-open-source-system-behind-better-harness" class="hash-link" aria-label="Direct link to The three-layer open-source system behind Better Harness" title="Direct link to The three-layer open-source system behind Better Harness" translate="no">​</a></h2>
<p>Publishing the <code>/better-harness</code> prompt on GitHub would technically qualify as
open source. But for a coding agent, no single prompt determines the result.
What matters is the full working method behind it: what is worth checking, what
counts as evidence, how judgments are formed, and how they can keep running and
changing in real projects.</p>
<p>That is why Better Harness opens three connected layers.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="layer-1-harness-engineering-best-practices">Layer 1: Harness Engineering best practices<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#layer-1-harness-engineering-best-practices" class="hash-link" aria-label="Direct link to Layer 1: Harness Engineering best practices" title="Direct link to Layer 1: Harness Engineering best practices" translate="no">​</a></h3>
<p>This layer answers a practical question: when we examine sessions, CLIs,
observability, Rules, Skills, MCP servers, memories, Hooks, and automation, what
should we check—and which conclusions cannot be drawn merely from the presence
of configuration?</p>
<p>The knowledge is organized by problem domain under <code>references/</code>:</p>
<table><thead><tr><th>Domain</th><th>Question it answers</th><th>Main evidence</th></tr></thead><tbody><tr><td><strong>Session Evidence</strong></td><td>How did the agent actually complete the task?</td><td>Sessions, task episodes, tool calls, retries, usage, and outcome evidence</td></tr><tr><td><strong>Project Harness</strong></td><td>Does the project provide reliable execution, verification, and delivery paths?</td><td>CLI, observability, design contracts, tests, Git Hooks, sensitive code, and recovery mechanisms</td></tr><tr><td><strong>Agent Customize</strong></td><td>Are agent assets discoverable, applicable, and actually useful?</td><td>Rules, Skills, MCP, memories, Hooks, Custom Agents, and host configuration</td></tr><tr><td><strong>Loop Engineering</strong></td><td>Which mechanism should own a confirmed repeated workflow?</td><td>Skills, Hooks, scripts, automation, Rules, Custom Agents, MCP, and related mechanisms</td></tr></tbody></table>
<p>Better Harness does not load one endlessly expanding master prompt on every
run. It reads the judgment criteria for the problem at hand. A diagnostic issue
routes to observability practices. A Skill issue routes to Skill Review. A
repeated workflow first triggers a decision about whether a Skill, Hook, script,
or automation should own it over time.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="layer-2-the-agent-work-loop-evaluation-model">Layer 2: the Agent Work Loop evaluation model<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#layer-2-the-agent-work-loop-evaluation-model" class="hash-link" aria-label="Direct link to Layer 2: the Agent Work Loop evaluation model" title="Direct link to Layer 2: the Agent Work Loop evaluation model" translate="no">​</a></h3>
<p>This layer turns engineering practices into questions that can be checked one
by one, while constraining the relationship between evidence, scores, and
conclusions. The current model and evidence states are public in the
<a href="https://github.com/QoderAI/better-harness/blob/main/models/agent-work-loop.md" target="_blank" rel="noopener noreferrer" class="">Agent Work Loop model</a>.</p>
<p>Because the standard is still emerging, we did not want one model to define a
“good harness” subjectively. The first internal evaluation selected 30 real
GitHub projects. Four model families independently evaluated them using
OpenAI's Harness Engineering article as a starting point, producing 120
standardized reports. Cross-model comparison and human calibration then
clarified evidence requirements, judgment boundaries, and differences between
project types before every project was evaluated again under the updated
criteria.</p>
<p>This loop—automated evaluation, automated aggregation, human calibration, and
automated reruns—produced the first reproducible Harness Engineering evaluation
model that we could continue to adjust.</p>
<p>The first version still resembled a conventional software-engineering maturity
scan. It focused on whether a project had documentation, tests, CI, and safety
mechanisms. We soon learned that static assets cannot prove that an agent
actually completed a task.</p>
<p>Better Harness itself is developed through a spec-driven process so that
changes in the model and product capabilities remain traceable. As more than
200 specs accumulated, the model shifted from asking “What exists in the
repository?” to asking “What actually happened in the task?”</p>
<p>The evaluation target narrowed from a repository or a session to a concrete
task. A session stopped being the thing being evaluated and became a container
for evidence. The model then stabilized around five dimensions: task
understanding, controlled execution, change verification, reliable delivery,
and experience capture. File existence, config counts, temporal proximity, and
even a successful command can no longer be treated as direct proof that a
capability was effective.</p>
<p>The Agent Work Loop is therefore not a static scorecard. It is a judgment
system centered on real tasks, designed to be reproduced and continuously
calibrated.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="layer-3-a-runnable-engineering-implementation">Layer 3: a runnable engineering implementation<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#layer-3-a-runnable-engineering-implementation" class="hash-link" aria-label="Direct link to Layer 3: a runnable engineering implementation" title="Direct link to Layer 3: a runnable engineering implementation" translate="no">​</a></h3>
<p>The third layer makes the practices and model repeatable in real projects.
Better Harness starts an analysis through a plugin or CLI—the JavaScript code
under the project's <code>scripts/</code> directory—freezes the task scope, and collects
three evidence lanes independently:</p>
<ul>
<li class=""><strong>Session Evidence</strong> reconstructs the agent's behavior in real tasks.</li>
<li class=""><strong>Project Harness</strong> checks whether the project can be started, diagnosed,
verified, and recovered.</li>
<li class=""><strong>Agent Customize</strong> checks the configuration, routing, and usage evidence for
Rules, Skills, MCP servers, memories, and Hooks.</li>
</ul>
<p>The lanes remain separate during collection and analysis. Only then does the
Lead reconcile them using the criteria in <code>references/</code> and the Agent Work Loop
model. “The project has this capability” and “the agent used this capability in
the task” remain two different facts.</p>
<p>The output is not just a score. It is a set of Findings with evidence
boundaries, user impact, repair scope, and verification methods. Once rendered
and validated, the report can enter a repair flow. If the analysis finds stable
repeated work, Loop Engineering determines whether a Skill, Hook, script,
automation, or another mechanism should own it over time.</p>
<p>Completing a repair still does not prove that the workflow improved. The loop
is closed only when a later task of the same kind produces a better observed
result.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="start-with-the-first-verifiable-problem">Start with the first verifiable problem<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#start-with-the-first-verifiable-problem" class="hash-link" aria-label="Direct link to Start with the first verifiable problem" title="Direct link to Start with the first verifiable problem" translate="no">​</a></h2>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="for-qoder-users">For Qoder users<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#for-qoder-users" class="hash-link" aria-label="Direct link to For Qoder users" title="Direct link to For Qoder users" translate="no">​</a></h3>
<p>At the time of the announcement, Qoder Desktop included Better Harness in the
Quest view as <strong>Better Harness (Beta)</strong> and exposed <code>/better-harness</code> directly.
Qoder CLI and the JetBrains plugin could use the same capability on a machine
where Qoder Desktop had already been installed. Refer to the current
<a href="https://qoderai.github.io/better-harness/docs/installation#qoder" target="_blank" rel="noopener noreferrer" class="">Installation guide</a>
for today's supported Qoder entrypoints.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="from-the-open-source-repository">From the open-source repository<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#from-the-open-source-repository" class="hash-link" aria-label="Direct link to From the open-source repository" title="Direct link to From the open-source repository" translate="no">​</a></h3>
<p>Visit the
<a href="https://github.com/QoderAI/better-harness" target="_blank" rel="noopener noreferrer" class="">Better Harness GitHub repository</a>
and follow the current README or Installation guide. For example, Claude Code
users can run:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">/plugin marketplace add QoderAI/better-harness</span><br></div><div class="token-line" style="color:#393A34"><span class="token plain">/plugin install better-harness@better-harness</span><br></div></code></pre></div></div>
<p>Then start a review with:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#393A34;--prism-background-color:#f6f8fa"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#393A34;background-color:#f6f8fa"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#393A34"><span class="token plain">/better-harness Analyze my project's harness and generate an HTML report</span><br></div></code></pre></div></div>
<p>Your first Better Harness run does not need to build a complete agent
engineering system, and it does not need to chase a perfect score. A more
practical starting point is one problem with clear evidence, a concrete impact,
and a fast verification path.</p>
<p>It might be a check command the agent cannot find, an error log with no useful
next diagnostic step, or a Skill that exists but has never entered the task
routing path.</p>
<p>Fix one problem, run the review again, and observe whether a later task of the
same kind changes. Harness Engineering is not a one-time configuration project.
It is the continuous work of making a project easier for an agent to understand,
execute, and verify.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="we-know-it-is-not-complete">We know it is not complete<a href="https://qoderai.github.io/better-harness/blog/better-harness-is-now-open-source#we-know-it-is-not-complete" class="hash-link" aria-label="Direct link to We know it is not complete" title="Direct link to We know it is not complete" translate="no">​</a></h2>
<p>Better Harness has been run and calibrated repeatedly in Qoder's real
development workflows, but the current model still reflects the project types
and task scenarios we know best.</p>
<p>Different technology stacks, project sizes, team constraints, and coding
agents may reveal blind spots. Better Harness needs more real evidence to keep
correcting them.</p>
<p>If you would like to contribute, there are several useful starting points:</p>
<ul>
<li class=""><strong>Add an engineering practice.</strong> Add a judgment guide for a language,
framework, or common workflow under <code>references/</code>. No code is required.</li>
<li class=""><strong>Add an evaluation perspective.</strong> Add an evidence-backed dimension or
detector under <code>models/</code> or <code>scripts/</code>, together with fixtures and tests.</li>
<li class=""><strong>Add host support.</strong> Complete evidence collection and verification for
another coding agent. The repository
<a href="https://github.com/QoderAI/better-harness/blob/main/roadmap.md" target="_blank" rel="noopener noreferrer" class="">Roadmap</a>
lists candidate work.</li>
<li class=""><strong>Add a real case study.</strong> Contribute a redacted team example under
<code>case-studies/</code>.</li>
</ul>
<p>If you disagree with a Finding, please open an issue. A counterexample from a
real project is more useful to us than a star—although we will happily accept
the star too. 😁</p>
<hr>
<p><strong>Born in Qoder, returned to the community. Give every coding agent a foundation
of verified engineering practice.</strong></p>]]></content:encoded>
            <category>better-harness</category>
            <category>open-source</category>
            <category>harness-engineering</category>
            <category>agent-work-loop</category>
        </item>
        <item>
            <title><![CDATA[Introducing Better Harness in Qoder]]></title>
            <link>https://qoderai.github.io/better-harness/blog/better-harness-in-qoder</link>
            <guid>https://qoderai.github.io/better-harness/blog/better-harness-in-qoder</guid>
            <pubDate>Thu, 30 Jul 2026 02:00:00 GMT</pubDate>
            <description><![CDATA[Learn how Better Harness diagnoses weak links in a coding-agent workflow, plans bounded improvements, and verifies whether the loop actually got better.]]></description>
            <content:encoded><![CDATA[<p>Today's coding agents can read requirements, modify code, run tests, and even
submit pull requests. But being able to do many things is not the same as being
able to do them well.</p>
<p>An agent usually cycles through <strong>understanding the task, taking action,
checking the result, and adjusting its next step</strong>. That is the Agent Loop. A
reliable loop does more than keep the agent moving: it gives the agent a clear
goal, defines what it must not touch, explains how to judge the result, and
provides a recovery path when something fails. Without those boundaries, an
agent may change a great deal of code and run many tests while still being
unable to prove that the task is actually complete.</p>
<p>This is the problem that Loop Engineering and Harness Engineering address.
They equip the agent with project context, relevant development tools,
effective verification methods, and explicit safety boundaries so that every
loop moves closer to a reliable delivery.</p>
<p>Building on Qoder's internal experience and the broader community's work on
coding agents, agent loops, and software engineering, we introduced <strong>Better
Harness (Beta)</strong>.</p>
<p>In current versions of Qoder, you can open Better Harness and start an analysis
and repair from the visual interface, or run the <code>/better-harness</code> Skill
directly. It examines how the agent worked through the current task, identifies
missing or weak elements in the loop, and helps you decide what to improve
next.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-better-harness-diagnoses-improves-and-rechecks-the-loop">How Better Harness diagnoses, improves, and rechecks the loop<a href="https://qoderai.github.io/better-harness/blog/better-harness-in-qoder#how-better-harness-diagnoses-improves-and-rechecks-the-loop" class="hash-link" aria-label="Direct link to How Better Harness diagnoses, improves, and rechecks the loop" title="Direct link to How Better Harness diagnoses, improves, and rechecks the loop" translate="no">​</a></h2>
<p>Better Harness does not grade the quality of a single answer. It examines the
entire harness that supports a coding agent as it completes a task:</p>
<ul>
<li class="">Are the goal and context clear?</li>
<li class="">Is the project easy to run?</li>
<li class="">Are permissions under control?</li>
<li class="">Does verification provide meaningful evidence?</li>
<li class="">Is delivery safe?</li>
<li class="">Can the team and the agent learn from the task?</li>
</ul>
<p>Its main analysis flow is:</p>
<ol>
<li class=""><strong>Map the current harness.</strong> Identify the goal, context, execution
entrypoints, feedback paths, delivery mechanisms, and learning mechanisms.</li>
<li class=""><strong>Find the breaks.</strong> Explain which part of the loop lacks a mechanism,
integration, observed execution, or outcome evidence.</li>
<li class=""><strong>Choose the smallest improvement vehicle.</strong> Route the problem to the most
appropriate Rule, Skill, Hook, script, automation, or human gate.</li>
<li class=""><strong>Repair and recheck.</strong> Keep the fix bounded, run the relevant verification,
and run <code>/better-harness</code> again to see whether the loop actually improved.</li>
</ol>
<p>The main analysis flow first collects the underlying evidence. It then asks
three independent, read-only subagents to interpret three evidence lanes:</p>
<ul>
<li class=""><strong>Agent customization assets</strong>, such as Rules, Skills, and Hooks;</li>
<li class=""><strong>Real task-session records</strong>, which show what the agent actually did and how
the task ended; and</li>
<li class=""><strong>The project's software-engineering foundation</strong>.</li>
</ul>
<p>The three lanes are collected independently and reconciled only afterward, so
that one category of evidence does not contaminate the conclusions drawn from
another.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="agent-customization-from-capability-inventory-to-actual-use">Agent customization: from capability inventory to actual use<a href="https://qoderai.github.io/better-harness/blog/better-harness-in-qoder#agent-customization-from-capability-inventory-to-actual-use" class="hash-link" aria-label="Direct link to Agent customization: from capability inventory to actual use" title="Direct link to Agent customization: from capability inventory to actual use" translate="no">​</a></h2>
<p>In coding agents such as Qoder, Rules, Skills, Custom Agents, MCP servers,
plugins, memories, and Hooks are the building blocks of an effective loop.
Better Harness inventories these custom capabilities and checks whether they
are complete, discoverable, and usable.</p>
<p>But having the building blocks does not make the loop reliable. A project can
have tests without running the relevant tests for a task. A Skill can exist
without the agent invoking it when it matters. Better Harness therefore also
examines evidence that Rules and Skills were actually used, helping users find
context-engineering gaps and reduce wasted credits.</p>
<p>Qoder Canvas presents this evidence in a detailed report, including patterns
such as Skill usage over a recent period. Better Harness also looks for
repeated work in task sessions that might justify a reusable Rule or Skill.
However, not every observation should become a Skill, and not every repeated
task should be automated. The report keeps those distinctions explicit and
offers bounded improvement suggestions.</p>
<p>For an identified opportunity, the user can select <strong>Plan a fix</strong> and let the
AI generate and execute a repair plan. More importantly, the result does not
have to remain a one-off fix. It can become a reusable Rule, Skill, memory, or
other asset that strengthens the user's own agent harness.</p>
<p>Each task analysis can therefore improve more than the current task. As the
asset base grows, the agent learns more about the user and the quality,
efficiency, and control of later loops can improve as well.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="real-task-sessions-reconstructing-how-the-agent-loop-actually-ran">Real task sessions: reconstructing how the Agent Loop actually ran<a href="https://qoderai.github.io/better-harness/blog/better-harness-in-qoder#real-task-sessions-reconstructing-how-the-agent-loop-actually-ran" class="hash-link" aria-label="Direct link to Real task sessions: reconstructing how the Agent Loop actually ran" title="Direct link to Real task sessions: reconstructing how the Agent Loop actually ran" translate="no">​</a></h2>
<p>Capability inventory alone cannot prove that a loop worked. Better Harness
also analyzes real task-session records—by default, from the most recent 30
days—to understand what the agent actually did and what outcome it produced.</p>
<p>The basic unit of analysis is a task episode: one user goal plus an observable
acceptance boundary. Within each episode, Better Harness looks for four kinds
of signals:</p>
<ol>
<li class=""><strong>Repeated workflows.</strong> When the same task repeatedly requires the same
steps or corrections, the project may be missing a Skill, Rule, or script.</li>
<li class=""><strong>Closed-loop verification.</strong> Did the agent actually run the relevant tests,
lint checks, builds, or regressions in the right place? Repeating a check is
not enough; the subsequent result must also be accepted and used.</li>
<li class=""><strong>Attribution of friction.</strong> When a task stalls, did the problem come from
the harness, the project, the model, or the requirement itself? Not every
failure should be blamed on the agent.</li>
<li class=""><strong>High-impact one-off events.</strong> Did a permission block, missing diagnostic
entrypoint, or failed recovery materially change the direction of the task?</li>
</ol>
<p>This lane is analyzed by an independent, read-only subagent. It sees only
redacted factual summaries, not raw prompts, private paths, secrets, or other
sensitive material. These signals make it possible to assess the real use of
Skills and Rules, identify context-engineering problems that affect the loop,
and reduce unnecessary credit consumption.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="project-engineering-foundations-make-it-findable-runnable-and-verifiable">Project engineering foundations: make it findable, runnable, and verifiable<a href="https://qoderai.github.io/better-harness/blog/better-harness-in-qoder#project-engineering-foundations-make-it-findable-runnable-and-verifiable" class="hash-link" aria-label="Direct link to Project engineering foundations: make it findable, runnable, and verifiable" title="Direct link to Project engineering foundations: make it findable, runnable, and verifiable" translate="no">​</a></h2>
<p>The project's existing software-engineering foundations also shape the quality
of every agent loop. A repository may contain extensive documentation, scripts,
and tests, but those capabilities cannot help if the agent cannot find the
correct entrypoint, run it successfully, or decide what to do after it fails.</p>
<p>Better Harness examines five aspects of a project's readiness for agent work:</p>
<ol>
<li class=""><strong>Findable.</strong> Can the agent quickly locate the relevant code, module
boundaries, project constraints, and required checks for a specific task?</li>
<li class=""><strong>Runnable.</strong> Are dependencies, configuration, build steps, and startup
instructions clear? When the environment breaks, can the agent diagnose it,
reset safely, and start again?</li>
<li class=""><strong>Fast feedback.</strong> After a code change, do the available checks and tests
quickly show what failed, where it may have failed, and what to try next?</li>
<li class=""><strong>Enforceable rules.</strong> Are architecture, security, API compatibility, and
database-migration requirements checked by tools rather than existing only
as documentation or convention?</li>
<li class=""><strong>Controlled changes.</strong> Are change boundaries explicit, do high-risk actions
require confirmation, and can a failed operation be rolled back, recovered,
or exited safely?</li>
</ol>
<p>For example, a test command in the README proves only that the project exposes
a test entrypoint. Inspecting the script reveals what the command actually
covers. Only running the relevant check in a real task and responding to its
result can show that the feedback path has truly entered the Agent Loop.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="turn-every-agent-loop-into-an-asset-for-the-next-one">Turn every Agent Loop into an asset for the next one<a href="https://qoderai.github.io/better-harness/blog/better-harness-in-qoder#turn-every-agent-loop-into-an-asset-for-the-next-one" class="hash-link" aria-label="Direct link to Turn every Agent Loop into an asset for the next one" title="Direct link to Turn every Agent Loop into an asset for the next one" translate="no">​</a></h2>
<p>A <code>/better-harness</code> analysis is not the finish line. It helps you find breaks,
generate a repair plan, and verify whether a new Rule, Skill, Hook, or script
actually entered the agent's workflow.</p>
<p>The most reusable lessons can then become personal or team-owned agent assets,
making later loops more stable, efficient, and controllable. Open Better
Harness in Qoder—or run <code>/better-harness</code>—and find the next part of your loop
that is worth strengthening.</p>]]></content:encoded>
            <category>better-harness</category>
            <category>qoder</category>
            <category>harness-engineering</category>
            <category>loop-engineering</category>
        </item>
    </channel>
</rss>