Weaver: An AI Self-Evolving System Grown from Practice
As a “poor student” who has heavily relied on geminicli for three months, while I am constantly amazed by Gemini’s intelligence, I am also frequently troubled by its forgetfulness and hallucinations.
The pain is visceral: you’re wrestling with complex generic logic in a large project. You clearly stated in the first turn, “All type definitions must use interface,” yet by the 30th turn, it’s casually writing type everywhere. Or it confidently reports, “Code modified successfully,” but when you open the file, the physical content remains untouched. Repeatedly correcting the same mistakes, or watching it mess up code in a large project while ignoring the facts—it’s a real headache.
Recently, openclaw went viral. I gave it a spin—it’s fun, but it didn’t quite feel like a reliable tool for research or development yet. However, the wisdom of the crowd is infinite. Browsing clawhub, I saw many brilliant skills. Inspired by the self-evolment approach, I looked through the geminicli extension library, found no similar plugins, and decided to build one myself. This was my first attempt at a ReAct system (Reasoning and Acting loop). My personal experience has been decent—memory has definitely improved—though long-term effects remain to be seen.
Looking back, I realized that only through daily practice can one uncover technical nuances and devise solutions. All progress is built on constant trial and error. As Marx said, all human knowledge comes from productive practice. Only by doing can we progress.
I asked the AI to summarize the process below. Feel free to read or try the plugin I developed (even if it only has one star from me right now).
🚀 Plugin Quick Start
Install Command:
gemini extensions install https://github.com/Biogod2020/ASSA
[!IMPORTANT] Current Limitations: The extension isn’t perfect. When used alongside other tools, the memory function sometimes fails to trigger automatically. Additionally, as the context grows longer, the hook injection seems to weaken.
In the future, I might consider more low-level logic to ensure memory is submitted consistently. But for now, with limited dev time, manually triggering it in long contexts works fine. I’m also waiting for Google to update and provide more robust, lower-level hooks. Feedback is always welcome!
01. Evolution Timeline: Eight Days of Fire
In just over a week, we underwent “compressed evolution”—from building vertical memory to restructuring horizontal order.
Day 1-2: Clueless Exploration
Trial and error with Python scripts, eventually migrating to the JS/TS ecosystem.
Day 3-4: Embracing MCP & Hooks
Solving instruction drift and execution black boxes; establishing the L1-L2-L3 distillation path.
Day 4: Solving the "Freshness" Problem
Introducing PENDING / PROCESSED status filtering to prevent historical amnesia.
Day 5: Architectural Foresight
Introducing **Graph** organization and categorizing knowledge into **G0-G3** tiers.
Day 6: Embracing Subagents
Offloading main process pressure; decoupling Distiller and Promoter logic.
Day 6-8: Peak Performance & Governance
Implementing the Index-First Strategy and establishing main process sovereignty.
02. Clueless Exploration (Day 1-2)
When development first began, I was truly wandering in the dark.
I initially asked the AI to write a version in Python, trying to intercept and record conversations with simple scripts. This immediately hit a wall: Python scripts required extra dependencies (like running pip install constantly), and they felt out of place within Gemini CLI’s native TypeScript/Node.js architecture. In practice, these cross-language calls made the environment unstable, leading to frequent freezes.
After two days of struggle, I abandoned Python and decided to follow Gemini CLI’s native ecosystem, switching entirely to JS/TS. This was the first step toward professional engineering—painful as it was to start over, it laid the foundation for the high-performance Hook mechanism that followed.
03. Embracing MCP & Hooks, Establishing the Evolution Path (Day 3-4)
After switching to JS, I began to study the tools Gemini CLI provides for developers.
I learned what “Hooks” are—they act as “undercover agents” inserted before AI reasoning (BeforeAgent) and after tool execution (AfterTool). I started using MCP (Model Context Protocol) tools to distill daily errors and corrections.
Through continuous discussion and experimentation, the AI and I co-summarized a highly effective concept: the L1-L2-L3 Knowledge Distillation Path.
- L1 (Ledger): Records raw correction signals and error messages like a ledger.
- L2 (Local): Distills these into development habits and patterns specific to the current project.
- L3 (Global): Promotes them to global guidelines across projects.
This hierarchical approach mirrors how human developers summarize their own experiences—knowledge is no longer a tangled mess, but a clear path for promotion.
Physical Metadata
We no longer rely solely on the AI's semantic memory. Instead, we use Hooks to forcefully inject physical identifiers into every tool output. This establishes objective coordinates, ensuring reported successes are based on actual file changes.
Semantic Reflex Sensor
Leveraging the AI's understanding of interaction sentiment. When I correct an error or offer praise (e.g., "Perfect," "That's wrong"), the system automatically captures these signals and triggers a feedback loop to record the lesson immediately.
04. Solving the Knowledge “Freshness” Problem (Day 4)
As interactions increased, it became clear that simply accumulating facts wouldn’t work. Log files grew longer, and if every distillation required reading the entire history, Token consumption and response times would skyrocket.
The AI suggested vector search or summary compression, but both felt too complex and prone to losing detail.
Instead, I came up with a “brute force” solution: classifying knowledge as either “Expired/Processed” or “Fresh.” I added status flags to every signal in the Ledger. This way, during distillation, the AI only processes “Fresh” knowledge marked as <span class="text-orange-500 font-bold text-sm">PENDING</span>, tagging it as <span class="text-green-500 font-bold text-sm">PROCESSED</span> once done. Efficiency surged, and the AI’s attention became laser-focused.
05. Architectural Foresight: Graph & G0-G3 Tiers (Day 5)
Even with efficient distillation, the rules in the files continued to pile up. Without a solid architecture from the start, organizing them would eventually become impossible.
Drawing on my experience with tools like Obsidian, I realized that a Graph structure would be ideal. I moved away from flat Markdown lists toward an interconnected knowledge graph.
I also felt that knowledge shouldn’t just be a flat graph—rules have different weights. For instance, not deleting code recklessly has higher priority than following a specific naming convention. This led to the G0-G3 tiered system:
06. Embracing Subagents (Day 6)
While the Graph worked well, asking the main process Agent to both write code and maintain a sophisticated knowledge graph system was a massive engineering load—context would blow up easily.
Looking for optimizations, I saw Subagent implementations in the Superpowers extension and realized Gemini CLI supports them natively.
I decided to embrace Subagents, offloading background distillation (Distiller) and global rule synchronization (Promoter) into independent sub-tools. This was like giving the main Agent two secretarial assistants, drastically reducing its workload while boosting overall system performance.
Division of Labor: Main Process vs. Subagents
Focuses on executing user tasks, writing core code, and making final Architectural Governance decisions.
Perform heavy-lifting in the background: transforming raw signals into Patterns and promoting them to the global library.
07. Extreme Optimization & Testing (Day 6 - Present)
Real-world testing and optimization are always the most tedious parts. To handle context bloat caused by injecting too many rules (which once hit 25KB+), I implemented an aggressive “Index-First” strategy, also known as Skeleton-First parsing.
The system no longer crams in the full text of every rule. Instead, it injects only the index skeleton, fostering a “pre-reading instinct” in the Agent so it can read_file on its own when it needs to make changes.
Furthermore, to prevent Subagents from making messy logical choices during organization, I established “Main Process Sovereignty”: Subagents handle the distillation and conflict detection, but the final decision on merging must be queried through the main process to the user.
After countless iterations, we arrived at version 3.5—finally, a system that’s truly a joy to use.
The System Codex: Unshakable Principles
During testing, I realized the system needed explicit “ground rules” to prevent the AI from getting lazy. I categorized these as G1 Engineering Standards and hardcoded them into its system prompt:
Conclusion: Knowledge Grown from Practice
Marx once said that all human knowledge comes from productive practice. Every functional node of ASSA wasn’t a pre-designed blueprint; it was “forged” by the AI and me through constant trial, anxiety, and correction.
The truth of engineering often hides within the most mundane failures. When you start taking every “historical misread” by the AI seriously, and when you start worrying that your knowledge base will become “messy,” the seeds of evolution are sown. The Weaver architecture isn’t a pre-set blueprint; it is a physical response to the pains of evolution.
The Philosophy of "Growing"
Initially, I wanted to design a perfect architecture from scratch, like a comprehensive Python framework. But experience proved that design decoupled from actual usage scenarios is just a fantasy. Structures that seem brilliant in theory collapse instantly when they hit real-world code errors, Token limits, or tool timeouts.
True iteration happens during every painful copy-paste, and every exasperated "Why did you forget again?!" It forces you to write a Hook, add a status flag, or split off a Subagent. This is the essence of "practice leads to true knowledge." Don't be afraid of messy code at the start. As long as you are writing, using, and feeling the pain, it will eventually grow into what it needs to be.
Through these three months of experimentation, I haven’t just built a better tool; I’ve deeply experienced the joy of moving forward through practice. While this plugin is still in its infancy, the process of watching it grow from nothing has been my greatest reward.