The Setup
So here’s the issue. Two weeks in a row I did the same thing.
Opened Substack. Copied last week’s notes into Claude. Said: “analyze these.” Got some generic output back. Closed the tab.
I kept telling myself the review mattered. And it did. But going through that process manually every Sunday was slow enough to make me frustrated every time.
The third Sunday I sat down to do it again, I stopped. Not to quit - to think.
What if I woke up the next Sunday morning with the review already done?
Unfortunately I had no way to pull the data automatically. No framework to assess against. No structured output to actually learn from.
So instead of fixing my habits, I fixed the system.
The System
Simple idea. Three inputs, one output:
Weekly Notes - every note I posted that week, with like count and timestamp
Framework Doc - a Notion page defining my content pillars, hook rules, writing style, and what good engagement actually signals
Weekly Report - one table: each note, the related stats (like count and comments), and what worked or didn’t
The framework doc isn’t 100% static. It updates when my thinking does - maybe once a month.
The Process
I spent a Sunday morning figuring out the pieces: what scraper to use, how to pull the framework doc, how to structure the final output.
Three stages ended up being enough.
The first uses Apify - a web scraping platform - to pull my recent notes automatically. Not just the text. The like count, the timestamp, everything I need without manually digging through the app.
The second uses Claude’s native Notion connector to retrieve the framework doc. No API key needed - just a one-time login. Claude reads the doc the same way it reads anything else you put in front of it.
The third is where it gets interesting. I built two lightweight custom skills inside Claude - one that handles the scraping, one that runs the review. Together they pass each note through the framework and ask the things I used to forget to ask:
Did the hook open with a real moment or story?
Is the angle specific or too broad?
Does the voice hold throughout the note or does it drift into generic advice?
The whole thing runs with a single command. One skill calls the other, pulls the framework from Notion, runs the review, and writes the output back - all without me touching anything.
Here’s how the pieces fit together:
This is how it looks in for the week of April 13 → April 18:
What It Caught That I Missed
The first week I ran it, it flagged three notes I’d written off. Not viral - just quietly above my usual numbers.
All three had one thing in common: they opened with something that happened to me, not something I concluded.
No lesson. No angle. Just a short story.
This is an example of such a post:
I also noticed I’d been going more opinionated lately because it felt bolder and punchier. I felt like the quiet ones were too simplistic and not worth sharing. But what the data actually cared about wasn’t how confident the opener was, but whether someone stopped scrolling or not.
One of those notes had no take at all. Something went sideways, I kept going anyway, no moral attached. It hit roughly 3x my average that week.
The system also flagged hooks that were too broad. Not bad - just unfocused. I’d been trying to expand the post under different angles, multiple topics sometimes, but the framework kept saying the same thing:
Pick one, go deep.
It’s annoying to hear, but that’s usually how you know something starts to work.
Engagement → Before vs After
Before the system: 4-5 likes per note. Occasionally 10-11 if something really landed.
After two weeks: 7-10 minimum, 10-15 on average. Posts that connect are hitting 25-35, but I treat those as outliers.
I have almost no audience on Substack. These aren’t numbers softened by thousands of followers. Each like is a real person who chose to stop and read the note. When that doubles in a small room, it’s not the algorithm. It’s the content.
The Part Nobody Mentions
This only works because I already had a framework.
If I hadn’t written out my content pillars, my hook rules, my style notes - the reviewer would have nothing to compare against. It’d return a word count and a shrug.
The automation is maybe 20% of this. The document it reads is the other 80.
Most people try to skip the document and go straight to the tool. They set up a scraper, ask Claude to review their posts, and wonder why the output feels hollow. And I’ll be honest, even with my current setup in 10% of cases it feels hollow.
It’s hollow because there’s nothing real to compare against.
Build the framework document first. Then automate around it.
The automation is nothing without something real to review against. Use AI to refine, not to generate the idea - that part still has to come from you.
If you want the full step-by-step to set this up, comment ‘notes’ below. First 10 people get it in their DMs.
Talk soon,
Ilya














now that growth feels real and smart..!!
Notes