<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[AIOps]]></title><description><![CDATA[AIOps]]></description><link>https://aiops-hub.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/69e7edc9e436727814a408ba/66904c29-73e0-4aa3-b49b-d3fe3dd198a3.jpg</url><title>AIOps</title><link>https://aiops-hub.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Thu, 08 Oct 2026 14:13:59 GMT</lastBuildDate><atom:link href="https://aiops-hub.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Building an AI-Powered CI/CD Auto-Triage and Healing Agent-Part-02]]></title><description><![CDATA[In Part 1, we built the simplest possible thing — a Python script that fetches a failed GitLab job trace and asks Claude to classify it. It worked. But the moment it worked, we saw every crack clearly]]></description><link>https://aiops-hub.hashnode.dev/building-an-ai-powered-ci-cd-auto-triage-and-healing-agent-part-02</link><guid isPermaLink="true">https://aiops-hub.hashnode.dev/building-an-ai-powered-ci-cd-auto-triage-and-healing-agent-part-02</guid><category><![CDATA[Devops]]></category><category><![CDATA[#AIOps]]></category><category><![CDATA[ai agents]]></category><category><![CDATA[automation]]></category><category><![CDATA[pydantic]]></category><category><![CDATA[langchain]]></category><dc:creator><![CDATA[Naveen Joshi]]></dc:creator><pubDate>Mon, 27 Apr 2026 13:16:51 GMT</pubDate><content:encoded><![CDATA[<p>In <a href="https://aiops-hub.hashnode.dev/building-an-ai-powered-ci-cd-auto-triage-and-healing-agent-part-01">Part 1</a>, we built the simplest possible thing — a Python script that fetches a failed GitLab job trace and asks Claude to classify it. It worked. But the moment it worked, we saw every crack clearly.</p>
<p>In this part we fix <strong>two of those cracks simultaneously</strong> because they are deeply connected:</p>
<ol>
<li><p><strong>CI logs are enormous</strong> — sending raw traces to an LLM is expensive, slow, and counterproductive</p>
</li>
<li><p><strong>LLM output is unstructured text</strong> — <code>json.loads()</code> on a Claude response is not a production strategy</p>
</li>
</ol>
<img src="https://cdn.hashnode.com/uploads/covers/69e7edc9e436727814a408ba/ea6b3334-a446-4e1a-8a42-fe811de7784b.jpg" alt="" style="display:block;margin:0 auto" />

<p>By the end of this post your app will process logs of any size efficiently and return a <strong>typed Python object</strong> you can actually work with downstream — not a string you hope parses correctly.</p>
<hr />
<h2>The Problem With Raw Traces</h2>
<p>Real CI/CD job traces are not the clean 50-line snippets in tutorials. A typical GitLab pipeline job produces:</p>
<pre><code class="language-plaintext">Docker layer pull logs     → 300-500 lines of layer hashes
npm/pip install output     → 200-400 lines of package resolution
Test setup and fixtures    → 100-200 lines of framework output
Your actual error          → 3-10 lines buried somewhere in the middle
</code></pre>
<p>A moderately complex pipeline job produces <strong>5,000 to 50,000 characters</strong> of output. Some jobs — particularly those running test suites or security scanners — produce <strong>100,000 characters or more</strong>.</p>
<p>Sending this raw to Claude has three real costs:</p>
<table>
<thead>
<tr>
<th>Problem</th>
<th>Impact</th>
</tr>
</thead>
<tbody><tr>
<td>Token cost</td>
<td>Every input token costs money. 50K chars ≈ 12K tokens ≈ $0.04 per job call</td>
</tr>
<tr>
<td>Quality degradation</td>
<td>LLMs struggle when the actual signal is buried in pages of noise</td>
</tr>
<tr>
<td>Latency</td>
<td>More tokens = longer inference time = slower Slack notifications</td>
</tr>
</tbody></table>
<p>The solution is not to hope Claude figures it out. It is to <strong>extract the signal before the LLM call</strong>.</p>
<hr />
<h2>Strategy 1 — Smart Truncation</h2>
<p>The simplest and most effective intervention. CI failure traces follow a predictable pattern: the error is almost always in the <strong>last 20-30% of the log</strong>. The first 70% is setup noise.</p>
<pre><code class="language-python">def prepare_trace_for_llm(trace: str, max_chars: int = 8000) -&gt; str:
    """
    Intelligently truncates a CI trace to fit context limits.
    Preserves the head (job context) and tail (where errors live).
    Drops the middle noise.
    """
    if len(trace) &lt;= max_chars:
        return trace

    # Keep a small head for job context, large tail for the error
    head_chars = 1000
    tail_chars = max_chars - head_chars

    head = trace[:head_chars]
    tail = trace[-tail_chars:]

    return f"{head}\n\n[... {len(trace) - max_chars} chars truncated ...]\n\n{tail}"
</code></pre>
<p>This one function cuts token cost by 60-80% for most jobs while preserving everything the LLM actually needs to classify the failure.</p>
<hr />
<h2>Strategy 2 — Error Signal Extraction</h2>
<p>Smart truncation handles size. Signal extraction handles quality. Instead of sending the full trace — even truncated — we pre-process it to pull out the lines that actually contain diagnostic information.</p>
<pre><code class="language-python">import re

def extract_error_signals(trace: str) -&gt; dict:
    """
    Extracts structured error signals from a raw CI trace.
    Returns a dict with the most diagnostic content isolated.
    This is what we pass to the LLM — not the raw trace.
    """
    lines = trace.split("\n")

    # Lines containing error markers
    error_lines = [
        line for line in lines
        if any(marker in line.lower() for marker in [
            "error", "exception", "failed", "fatal",
            "traceback", "exit code", "killed", "timeout",
            "not found", "permission denied", "cannot"
        ])
    ]

    # Last 100 lines — where failures always surface
    tail_lines = lines[-100:]

    # Python traceback if present
    traceback = None
    tb_match = re.search(
        r"Traceback \(most recent call last\):.*?(\w+Error[^\n]*)",
        trace,
        re.DOTALL
    )
    if tb_match:
        traceback = tb_match.group(0)[-2000:]

    # Exit code
    exit_code = "unknown"
    exit_match = re.search(r"exit(?:ed)? (?:with )?(?:code )?(\d+)", trace, re.I)
    if exit_match:
        exit_code = exit_match.group(1)

    return {
        "exit_code":   exit_code,
        "error_lines": "\n".join(error_lines[-50:]),   # last 50 error lines
        "tail":        "\n".join(tail_lines),
        "traceback":   traceback or "None detected",
    }
</code></pre>
<p>The <code>error_lines</code> field alone typically reduces a 50,000-character trace to <strong>500-800 characters</strong> of pure signal. The LLM now has exactly what it needs to classify the failure — nothing more.</p>
<hr />
<h2>Strategy 3 — Two-Pass Summarisation for Massive Traces</h2>
<p>For traces over 100,000 characters (security scanners, full test suite runs), even smart truncation may not be enough. Use a fast cheap model to summarise first, then pass the summary to your main model.</p>
<pre><code class="language-python">import boto3, json

bedrock = boto3.client("bedrock-runtime", region_name="us-east-1")

def summarise_large_trace(trace: str) -&gt; str:
    """
    Uses Claude Haiku (fast, cheap) to compress a massive trace
    into a diagnostic summary before the main analysis call.
    Reduces a 100K char trace to ~500 chars of signal.
    """
    prompt = f"""Summarise this CI/CD job log in under 400 words.
Focus ONLY on: what failed, the exact error message, what line or file caused it.
Ignore all successful steps. Be precise.

LOG:
{trace[-15000:]}"""

    response = bedrock.invoke_model(
        modelId="anthropic.claude-haiku-4-5-20251001",
        body=json.dumps({
            "anthropic_version": "bedrock-2023-05-31",
            "max_tokens":        600,
            "messages": [{"role": "user", "content": prompt}]
        })
    )
    body = json.loads(response["body"].read())
    return body["content"][0]["text"]


def prepare_trace_input(trace: str, job_name: str) -&gt; str:
    """
    Routing function — picks the right preparation strategy
    based on trace size. Returns a clean string ready for the LLM.
    """
    signals = extract_error_signals(trace)

    if len(trace) &gt; 100_000:
        # Two-pass for massive traces
        summary = summarise_large_trace(trace)
        return f"SUMMARY:\n{summary}\n\nKEY SIGNALS:\n{signals['error_lines']}"

    elif len(trace) &gt; 8_000:
        # Extract signals + tail for medium traces
        return (
            f"EXIT CODE: {signals['exit_code']}\n\n"
            f"ERROR LINES:\n{signals['error_lines']}\n\n"
            f"TRACEBACK:\n{signals['traceback']}\n\n"
            f"LOG TAIL:\n{signals['tail']}"
        )

    else:
        # Small trace — send as-is
        return trace
</code></pre>
<p><strong>Practical size guide:</strong></p>
<table>
<thead>
<tr>
<th>Trace size</th>
<th>Strategy</th>
<th>Typical token cost</th>
</tr>
</thead>
<tbody><tr>
<td>&lt; 8K chars</td>
<td>Send as-is</td>
<td>~2K tokens</td>
</tr>
<tr>
<td>8K–100K chars</td>
<td>Signal extraction + tail</td>
<td>~800 tokens</td>
</tr>
<tr>
<td>&gt; 100K chars</td>
<td>Two-pass Haiku → Sonnet</td>
<td>~400 tokens</td>
</tr>
</tbody></table>
<hr />
<h2>The Second Problem — Unreliable Output</h2>
<p>With signal extraction solved, we now address output reliability. In Part 1 we asked Claude to return JSON and used <code>json.loads()</code> on the response. This works most of the time.</p>
<p><em>Most of the time</em> is not acceptable in a system that triggers merge requests.</p>
<p>The failure modes are real:</p>
<pre><code class="language-python"># Claude responds with a preamble
"Here is the JSON you requested:\n```json\n{\"type\": \"dependency\"}"
# json.loads() → fails

# Claude uses slightly different key names
{"failure_type": "dependency"}  # you expected "type"
# KeyError downstream

# Claude adds an extra field you didn't ask for
{"type": "dependency", "confidence": "high", "suggested_action": "..."}
# Your downstream code breaks on unexpected fields
</code></pre>
<p>The solution is <strong>Pydantic with LangChain's output parser</strong>. The parser handles JSON extraction, validation, and instantiation into a typed Python object in one step.</p>
<hr />
<h2>Introducing LangChain LCEL and PydanticOutputParser</h2>
<p>LangChain's <strong>LCEL (LangChain Expression Language)</strong> is a composable way to build LLM pipelines using the <code>|</code> pipe operator. Each component in the chain has a defined input type and output type, and they connect only when those types are compatible.</p>
<p>The pattern we are building is:</p>
<pre><code class="language-plaintext">prompt_template | llm | parser
</code></pre>
<p>Where each <code>|</code> represents a type-safe handoff:</p>
<pre><code class="language-plaintext">prompt_template  →  outputs ChatPromptValue (formatted messages)
llm              →  consumes ChatPromptValue, outputs AIMessage
parser           →  consumes AIMessage, outputs your Pydantic object
</code></pre>
<p>Let's build it step by step.</p>
<hr />
<h2>Step 1 — Define Your Output Schema With Pydantic</h2>
<pre><code class="language-python">from pydantic import BaseModel, Field
from typing import Literal

class JobTriage(BaseModel):
    """
    Structured output from the triage LLM call.
    Every field is typed and validated — no dict key surprises.
    """
    status: str = Field(
        description="Always 'failed' for failed jobs"
    )
    failure_category: str = Field(
        description=(
            "Category: dependency | syntax | test | "
            "config | infra | security | flaky | unknown"
        )
    )
    failed_reason: str = Field(
        description="One clear sentence explaining what failed and why"
    )
    suggested_fix: str = Field(
        description="One actionable sentence describing how to fix it"
    )
    confidence: Literal["high", "medium", "low"] = Field(
        description="Confidence in the classification"
    )
</code></pre>
<p>Pydantic gives you:</p>
<ul>
<li><p><strong>Type validation</strong> — <code>confidence</code> can only be <code>"high"</code>, <code>"medium"</code>, or <code>"low"</code></p>
</li>
<li><p><strong>Attribute access</strong> — <code>result.confidence</code> instead of <code>result["confidence"]</code></p>
</li>
<li><p><strong>IDE autocompletion</strong> — your editor knows the shape of the object</p>
</li>
<li><p><strong>Downstream type safety</strong> — passing <code>JobTriage</code> objects around is self-documenting</p>
</li>
</ul>
<hr />
<h2>Step 2 — The PydanticOutputParser</h2>
<pre><code class="language-python">from langchain.output_parsers import PydanticOutputParser

parser = PydanticOutputParser(pydantic_object=JobTriage)

# This generates a JSON schema instruction string
# that tells Claude exactly what to return
format_instructions = parser.get_format_instructions()
</code></pre>
<p><code>get_format_instructions()</code> produces a JSON schema string that gets injected into your prompt. Claude reads it and knows precisely what keys, types, and values to return. The LLM is constrained by schema — not just by a casual "return JSON" request.</p>
<hr />
<h2>Step 3 — The Prompt Template</h2>
<pre><code class="language-python">from langchain_core.prompts import ChatPromptTemplate

SYSTEM_PROMPT = """You are a DevOps expert analysing CI/CD pipeline failures.
Analyse the provided CI log signals and classify the failure.
{format_instructions}"""

prompt_template = ChatPromptTemplate.from_messages([
    ("system", SYSTEM_PROMPT),
    ("user",   "Job: {job_name}\n\nLog signals:\n{log_trace_signal}")
])

# Bind format_instructions as a partial — it never changes between calls
# so we do not pass it at invoke() time
prompt_template = prompt_template(
    format_instructions=format_instructions
)
</code></pre>
<h2>Step 4 — The LLM</h2>
<pre><code class="language-python">from langchain_aws import ChatBedrock

llm = ChatBedrock(
    model_id="anthropic.claude-3-5-sonnet-20241022-v2:0",
    region_name="us-east-1"
)
</code></pre>
<hr />
<h2>Step 5 — Chain Everything Together</h2>
<pre><code class="language-python"># The chain — three components connected by pipe operators
chain = prompt_template | llm | parser
</code></pre>
<p>That is the entire chain. Three lines. Here is what each <code>|</code> does:</p>
<pre><code class="language-plaintext">prompt_template
  Receives:  {"job_name": "...", "log_trace_signal": "..."}
  Produces:  ChatPromptValue (system + user messages formatted)
       |
      llm
  Receives:  ChatPromptValue
  Produces:  AIMessage (raw response with JSON string in .content)
       |
    parser
  Receives:  AIMessage
  Produces:  JobTriage (validated Pydantic object)
</code></pre>
<p>After the <code>parser</code> step, <code>json.loads()</code> failures, key errors, and unexpected fields are all handled automatically.</p>
<hr />
<h2>Step 6 — Invoke the Chain</h2>
<pre><code class="language-python">def analyse_failed_job(job_name: str, trace: str) -&gt; JobTriage:
    """
    Full pipeline: extract signals → invoke chain → return typed object.
    """
    # Signal extraction — replaces sending raw trace
    log_trace_signal = prepare_trace_input(trace, job_name)

    # Chain invocation — returns a JobTriage object, not a dict
    result: JobTriage = chain.invoke({
        "job_name":         job_name,
        "log_trace_signal": log_trace_signal
    })

    return result
</code></pre>
<hr />
<h2>The Full Updated main.py</h2>
<pre><code class="language-python">from gitlab_client import get_failed_jobs, get_job_trace
from triage import analyse_failed_job
import os

def main():
    pipeline_id = os.environ["FAILED_PIPELINE_ID"]
    print(f"Analysing pipeline {pipeline_id}...")

    failed_jobs = get_failed_jobs(pipeline_id)
    if not failed_jobs:
        print("No failed jobs found.")
        return

    for job in failed_jobs:
        job_id   = str(job["id"])
        job_name = job["name"]

        trace  = get_job_trace(job_id)
        result = analyse_failed_job(job_name, trace)

        # result is now a JobTriage object — typed, validated
        print(f"\n--- {job_name} ---")
        print(f"Category:   {result.failure_category}")
        print(f"Reason:     {result.failed_reason}")
        print(f"Fix:        {result.suggested_fix}")
        print(f"Confidence: {result.confidence}")

if __name__ == "__main__":
    main()
</code></pre>
<hr />
<h2>What the Output Looks Like Now</h2>
<pre><code class="language-plaintext">Analysing pipeline 882341...

--- test:unit ---
Category:   dependency
Reason:     ModuleNotFoundError for 'fastavro' — not in requirements.txt
Fix:        Add fastavro&gt;=1.7.0 to requirements.txt
Confidence: high

--- build:docker ---
Category:   config
Reason:     Base image python:3.11-alpine not found — tag changed upstream
Fix:        Update Dockerfile FROM to python:3.11-alpine3.18
Confidence: high
</code></pre>
<p>Same useful output as Part 1 — but now:</p>
<ul>
<li><p>Typed <code>JobTriage</code> objects instead of raw dicts</p>
</li>
<li><p><code>result.confidence</code> never throws a <code>KeyError</code></p>
</li>
<li><p><code>Literal["high", "medium", "low"]</code> validation means the value is always one of three strings</p>
</li>
<li><p>Token cost reduced by 70% on typical traces</p>
</li>
<li><p>Parser handles any JSON formatting quirks Claude introduces</p>
</li>
</ul>
<hr />
<h2>Understanding the LCEL Type Contract</h2>
<p>One thing worth understanding clearly before moving forward: the <code>|</code> pipe operator in LCEL is not just aesthetic — it enforces <strong>type contracts at connection time</strong>.</p>
<pre><code class="language-plaintext">Works:
  prompt_template | llm        → prompt outputs messages, llm accepts messages ✅
  llm | parser                 → llm outputs AIMessage, parser accepts AIMessage ✅

Fails at runtime:
  parser | llm                 → parser outputs Pydantic object, llm expects messages ❌
  prompt_template | parser     → prompt outputs messages, parser expects AIMessage ❌
</code></pre>
<p>LangChain does not enforce these contracts at definition time — it fails at invocation. This means: <strong>define your chain once and test it immediately with a real call.</strong> If the types are wrong you will know straight away.</p>
<p>The useful mental model is that each component is a function with a specific signature, and <code>|</code> is function composition. You would not write <code>parser(prompt_template(input))</code> in raw Python — and LCEL is the clean way to express the same idea.</p>
<hr />
<hr />
<h2>What We Have Now vs Part 1</h2>
<table>
<thead>
<tr>
<th>Capability</th>
<th>Part 1</th>
<th>Part 2</th>
</tr>
</thead>
<tbody><tr>
<td>Input size</td>
<td>Full raw trace (expensive)</td>
<td>Extracted signals (efficient)</td>
</tr>
<tr>
<td>Token cost per job</td>
<td>~12K tokens</td>
<td>~800 tokens</td>
</tr>
<tr>
<td>Output type</td>
<td><code>dict</code> from <code>json.loads()</code></td>
<td>Typed <code>JobTriage</code> object</td>
</tr>
<tr>
<td>Output validation</td>
<td>None</td>
<td>Pydantic schema enforcement</td>
</tr>
<tr>
<td>Failure modes</td>
<td><code>json.loads()</code> errors, <code>KeyError</code></td>
<td>Handled by parser automatically</td>
</tr>
<tr>
<td>IDE support</td>
<td>None — raw dict</td>
<td>Full autocomplete on typed fields</td>
</tr>
</tbody></table>
<hr />
<h2>What's Still Missing</h2>
<p>With efficient input and structured output working, the system classifies failures accurately and cheaply. But every classification is based on <strong>general LLM knowledge</strong> — what Claude learned during training.</p>
<p>It does not know:</p>
<ul>
<li><p>Your organisation's internal packages and conventions</p>
</li>
<li><p>Your team's known fix patterns for recurring failures</p>
</li>
<li><p>Past successful fixes your team has applied to similar errors</p>
</li>
</ul>
<p>In Part 3 we add this contextual knowledge using <strong>Retrieval-Augmented Generation (RAG)</strong> with AWS Bedrock Knowledge Base — so the LLM's analysis is grounded in your organisation's own history, not just the internet.</p>
]]></content:encoded></item><item><title><![CDATA[Building an AI-Powered CI/CD Auto-Triage and Healing Agent-Part-01]]></title><description><![CDATA[Every failed pipeline is a detective story buried in thousands of lines of output. Here's how we built the first version of an agent that reads them so you don't have to.


The Problem We Are Tired Of]]></description><link>https://aiops-hub.hashnode.dev/building-an-ai-powered-ci-cd-auto-triage-and-healing-agent-part-01</link><guid isPermaLink="true">https://aiops-hub.hashnode.dev/building-an-ai-powered-ci-cd-auto-triage-and-healing-agent-part-01</guid><category><![CDATA[Devops]]></category><category><![CDATA[#AIOps]]></category><category><![CDATA[ai agents]]></category><category><![CDATA[langchain]]></category><dc:creator><![CDATA[Naveen Joshi]]></dc:creator><pubDate>Mon, 27 Apr 2026 12:26:14 GMT</pubDate><content:encoded><![CDATA[<blockquote>
<p><em>Every failed pipeline is a detective story buried in thousands of lines of output. Here's how we built the first version of an agent that reads them so you don't have to.</em></p>
</blockquote>
<hr />
<h2>The Problem We Are Tired Of</h2>
<p>It starts with a Slack message. "Pipeline failed." You open GitLab, find the failed job, scroll through hundreds of lines of output — Docker pulling layers, npm install noise, test setup logs — to find the three lines that actually matter. Then you fix it, push, wait, and do it again.</p>
<p>For a busy team running dozens of pipelines a day, this is not a minor inconvenience. It is a significant drain on engineering time. The worst part? Most failures repeat. The same dependency conflict. The same schema mismatch. The same environment variable that someone forgot to set. An engineer should not need to read the same error twice.</p>
<blockquote>
<p><em>The idea is simple: if an LLM can understand a Python/Node error from a Stack Overflow post, it can surely understand a CI log. Let's find out</em></p>
</blockquote>
<h2><strong>What We're Building in This Series</strong></h2>
<p>This is Part 1 of a series documenting a real project built over several months. We start with a single, humble LLM call and end with a production multi-agent system that automatically detects, classifies, and fixes CI/CD pipeline failures — creating merge requests with the corrected code before a human has to get involved.</p>
<img alt="" />

<p><strong>The Tech Stack</strong></p>
<img src="https://cdn.hashnode.com/uploads/covers/69e7edc9e436727814a408ba/df8b8ad1-bc30-4d4a-8e10-401b81c91433.png" alt="" style="display:block;margin:0 auto" />

<p><strong>Why Bedrock?</strong></p>
<p>Bedrock allow IAM-based auth rather than managing API keys. Bedrock gives you access to Claude, Titan, and others through the same AWS credentials your CI runner already has. No new secrets to manage.</p>
<h2><strong>The First Working Version</strong></h2>
<img src="https://cdn.hashnode.com/uploads/covers/69e7edc9e436727814a408ba/52a82593-5505-4fe9-8ce3-7b716ca392ed.jpg" alt="" style="display:block;margin:0 auto" />

<p>The simplest possible thing that could work: fetch the trace of a failed GitLab job, send it to Claude with a prompt asking for a classification, and print the result. No framework. No orchestration. Just Python, boto3, and a prompt.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69e7edc9e436727814a408ba/8c3565ea-2adb-4d6f-9584-944e7bfb58b3.png" alt="" style="display:block;margin:0 auto" />

<h3><strong>Step 1 — Fetch the Failed Job Trace from GitLab</strong></h3>
<p>GitLab exposes a <code>/jobs/{id}/trace</code> endpoint that returns the raw console output of any CI job. With a personal access token and the project ID, we can pull this programmatically.</p>
<pre><code class="language-python">import requests
import os

GITLAB_URL   = "https://gitlab.com"
GITLAB_TOKEN = os.environ["GITLAB_TOKEN"]  # personal access token
PROJECT_ID   = os.environ["CI_PROJECT_ID"]   # available in GitLab CI env

def get_failed_jobs(pipeline_id: str) -&gt; list:
    """Return all failed jobs for a given pipeline."""
    headers = {"PRIVATE-TOKEN": GITLAB_TOKEN}
    url = f"{GITLAB_URL}/api/v4/projects/{PROJECT_ID}/pipelines/{pipeline_id}/jobs"

    response = requests.get(url, headers=headers)
    response.raise_for_status()

    # Return only the jobs that actually failed
    return [job for job in response.json() if job["status"] == "failed"]


def get_job_trace(job_id: str) -&gt; str:
    """Fetch the raw console trace for a specific job."""
    headers = {"PRIVATE-TOKEN": GITLAB_TOKEN}
    url = f"{GITLAB_URL}/api/v4/projects/{PROJECT_ID}/jobs/{job_id}/trace"

    response = requests.get(url, headers=headers)
    response.raise_for_status()

    return response.text  # raw log output — can be 50K+ characters
</code></pre>
<h3><strong>Step 2 — Call the LLM With the Trace</strong></h3>
<p>AWS Bedrock exposes Claude via the <code>bedrock-runtime</code> boto3 client. You construct a message payload in Anthropic's format, call <code>invoke_model()</code>, and parse the response. That's it.</p>
<pre><code class="language-python">import boto3
import json

# bedrock-runtime — the client for invoking LLMs
# Note: this is different from "bedrock-agent-runtime" (Knowledge Base client)
bedrock = boto3.client("bedrock-runtime", region_name="us-east-1")

MODEL_ID = "anthropic.claude-3-5-sonnet-20241022-v2:0"

def classify_failure(job_name: str, trace: str) -&gt; dict:
    """
    Send a job trace to Claude and ask for a failure classification.
    Returns a dict with type, reason, suggested fix, and confidence.
    """

    prompt = f"""You are a DevOps engineer analysing a failed CI/CD job.

Job name: {job_name}

CI log:
{trace}

Analyse the failure and respond with a JSON object containing:
{{
  "failure_type": "one of: dependency | syntax | test | config | infra | flaky",
  "failed_reason": "one clear sentence explaining what failed and why",
  "suggested_fix": "one actionable sentence describing how to fix it",
  "confidence": "high | medium | low"
}}

Respond with valid JSON only. No preamble."""

    response = bedrock.invoke_model(
        modelId=MODEL_ID,
        body=json.dumps({
            "anthropic_version": "bedrock-2023-05-31",
            "max_tokens": 512,
            "messages": [{
                "role":    "user",
                "content": prompt
            }]
        })
    )

    # Parse the response body
    body    = json.loads(response["body"].read())
    raw_txt = body["content"][0]["text"]

    return json.loads(raw_txt)
</code></pre>
<h3><strong>Step 3 — Wire It Together</strong></h3>
<p>The entry point loops over every failed job in a pipeline, classifies each one, and prints a summary. At this stage we're running this as a downstream GitLab CI job that triggers when any pipeline in the group fails.</p>
<pre><code class="language-python">from gitlab_client import get_failed_jobs, get_job_trace
from triage import classify_failure
import os, json

def main():
    pipeline_id = os.environ["FAILED_PIPELINE_ID"]  # passed via trigger variables

    print(f"🔍 Analysing pipeline {pipeline_id}...")
    failed_jobs = get_failed_jobs(pipeline_id)

    if not failed_jobs:
        print("No failed jobs found.")
        return

    print(f"Found {len(failed_jobs)} failed job(s)\n")

    for job in failed_jobs:
        job_id   = job["id"]
        job_name = job["name"]

        print(f"--- Job: {job_name} ({job_id}) ---")

        trace  = get_job_trace(job_id)
        result = classify_failure(job_name, trace)

        print(f"Type:      {result['failure_type']}")
        print(f"Reason:    {result['failed_reason']}")
        print(f"Fix:       {result['suggested_fix']}")
        print(f"Confidence:{result['confidence']}\n")

if __name__ == "__main__":
    main()
</code></pre>
<h3><strong>What It Looks Like Running</strong></h3>
<pre><code class="language-shell">$ python main.py

🔍 Analysing pipeline 882341...
Found 2 failed job(s)

--- Job: test:unit (job_id: 4421098) ---
Type:       dependency
Reason:     ModuleNotFoundError for 'fastavro' — package missing from requirements.txt
Fix:        Add fastavro to requirements.txt and re-run the pipeline
Confidence: high

--- Job: build:docker (job_id: 4421099) ---
Type:       config
Reason:     Base image 'python:3.11-alpine' not found — image tag changed upstream
Fix:        Update Dockerfile FROM to python:3.11-alpine3.18 or use python:3.11-slim
Confidence: high
</code></pre>
<p>Alternatively , you can publish the triage for a pipeline failure to a slack channel</p>
<p>It works. Two jobs, two accurate classifications, two actionable fixes — all from raw log output that an engineer would have spent 15 minutes reading. The entire execution takes about 8 seconds.</p>
<h2><strong>The Reality Check — What's Already Broken</strong></h2>
<p>The moment it works is also the moment you see every limitation clearly. Here's the honest assessment after running it against a real day of pipeline failures:</p>
<ol>
<li><p><strong>Problem 1 — Context Window</strong></p>
<p>Real CI traces are often 50,000–200,000 characters. We're sending the entire raw trace to Claude. This is expensive, sometimes hits limits, and buries the actual error in noise. We fix this in next part</p>
</li>
<li><p><strong>Problem 2 — Unstructured Output</strong></p>
<p>We asked the LLM to return JSON — and it mostly does. But sometimes it adds a preamble. Sometimes the keys are slightly different <code>json.loads()</code> fails silently. The fix — proper structured output with Pydantic — comes in subsequent parts.</p>
</li>
<li><p><strong>Problem 3 — No Org Context</strong></p>
<p>The LLM gives you generic internet fixes. It doesn't know your internal packages, your team's conventions, or your past fixes for the same errors. RAG comes in subsequent parts.</p>
</li>
<li><p><strong>Problem 4 — One Job at a Time</strong></p>
<p>Multiple failed jobs in a pipeline produce multiple separate outputs with no consolidated view. Each is analysed in isolation so the LLM can't detect that three failures share one root cause. We redesign this in subsequent parts.</p>
</li>
</ol>
]]></content:encoded></item></channel></rss>