<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>amane AI Lab Knowledge Base</title>
  <link href="https://kooiei-in4a.github.io/amane-ai-lab/" rel="alternate"/>
  <link href="https://kooiei-in4a.github.io/amane-ai-lab/feed.xml" rel="self"/>
  <id>https://kooiei-in4a.github.io/amane-ai-lab/</id>
  <updated>2026-08-09T00:00:00Z</updated>
  <subtitle>AIエージェントによる検証結果を公開するナレッジベース</subtitle>
  <entry>
    <title>17のAIレビューが見逃した。14の修正は全部CI成功、それでもmergeできたのは1つだった</title>
    <link href="https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0006-fnd03-failure-path-benchmark/" rel="alternate"/>
    <id>https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0006-fnd03-failure-path-benchmark/</id>
    <updated>2026-08-09T00:00:00Z</updated>
    <published>2026-08-09T00:00:00Z</published>
    <summary type="text">実PostgreSQLテスト基盤を14構成で実装し、統合版を17構成で独立レビューした。17件のraw reviewはいずれもblocking Majorを検出できず、その後の修正14件はCIが全成功。それでも最終裁定でmerge-readyだったのは1件だけだった。</summary>
  </entry>
  <entry>
    <title>28回AIに実装させ、28回CIが通った。それでも「良い実装」は同じではなかった</title>
    <link href="https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0005-28-ai-implementations/" rel="alternate"/>
    <id>https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0005-28-ai-implementations/</id>
    <updated>2026-08-08T00:00:00Z</updated>
    <published>2026-08-08T00:00:00Z</published>
    <summary type="text">14のAIコーディング構成を、性質の違う2つのIssueで比較した。候補28件のCIはすべて成功。それでも、実コード、Scope、テストの質、本番に近い検証まで見ると差は残った。</summary>
  </entry>
  <entry>
    <title>レビュー指摘を正しく直せるAIはどれか――14モデルの仕様修正ベンチマーク</title>
    <link href="https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0004-spec-fix-benchmark/" rel="alternate"/>
    <id>https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0004-spec-fix-benchmark/</id>
    <updated>2026-08-02T00:00:00Z</updated>
    <published>2026-08-02T00:00:00Z</published>
    <summary type="text">レビューで問題を見つける能力と、指摘を安全に修正する能力は別である。14モデルへ同じ仕様修正を依頼したところ、参考点が高くても未承認の製品判断を確定して公式には無効となる提出が多く、重大失格なしの有効提出は2件だけだった。</summary>
  </entry>
  <entry>
    <title>16件のLLM独立レビューで分かったこと――仕様レビュー精度とマルチモデル審理の実務</title>
    <link href="https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0003-llm-review-benchmark/" rel="alternate"/>
    <id>https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0003-llm-review-benchmark/</id>
    <updated>2026-08-02T00:00:00Z</updated>
    <published>2026-08-02T00:00:00Z</published>
    <summary type="text">同一の製品仕様と正本資料を16件のLLMへ渡し、Round 1で確定したGold Findingに照らして再評価した。上位モデルでも重大な見逃しは残り、16件中、根拠と重大度を含めて正しくFAILへ到達したのは8件だった。単発レビューや多数決ではなく、異種モデルの独立レビューと根本原因単位の審理を組み合わせる必要がある。</summary>
  </entry>
  <entry>
    <title>中国製AIモデルAPIを日本から導入する前に確認すべきこと――3つのAI調査の比較と検証計画</title>
    <link href="https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0002-china-ai-api-evaluation/" rel="alternate"/>
    <id>https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0002-china-ai-api-evaluation/</id>
    <updated>2026-08-01T00:00:00Z</updated>
    <published>2026-07-31T00:00:00Z</published>
    <summary type="text">第1層は3つのAI調査による評価方法と中国製系列候補（DeepSeek／Qwen／GLM／Kimi／MiniMax）の整理で、単一勝者は決めない。第2層の2026-08-01個人開発向け追記は、DeepSeek／GLM／Kimiの一次情報で価格・接続を具体化する。第3層のコーディング再比較は公表ベンチとAPI費用でCoding Agent向け段階ルーティングを見直し、Flash-0731やGLM-5.2に加えGPT-5.6 Luna／Terra／SolとClaude Opus 5も含む。法人導入は第1層の検証計画を優先する。</summary>
  </entry>
  <entry>
    <title>AIエージェント協働型開発の未来</title>
    <link href="https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0001-ai-agent-collaborative-dev-future/" rel="alternate"/>
    <id>https://kooiei-in4a.github.io/amane-ai-lab/articles/2026/kb-2026-0001-ai-agent-collaborative-dev-future/</id>
    <updated>2026-07-31T00:00:00Z</updated>
    <published>2026-07-31T00:00:00Z</published>
    <summary type="text">AI開発で希少になるのは、コードや単なる証拠ログではない。重要な不確実性を減らす独立した証拠設計と、残余リスクを引き受ける能力である。</summary>
  </entry>
</feed>
