SMF·WRITERZEN
Can a prompt replace WriterZen?
AI writing — AI drafting, research and content optimization
Exhibit tracking slip
Verdict
WriterZen's clustering is not semantic — it groups two keywords together when their top-ten results share enough of the same URLs, which is the only definition that reflects how a search engine treats them. That method is entirely reproducible and needs no model at all. The cost is that every keyword in the list needs a SERP fetched before it can be clustered, which is a paid API call each, and the bill scales with exactly the input size that makes clustering worth doing.
Exhibit A — The prompt
Received on31.07.2026Build a keyword clustering tool using SERP overlap, not semantic similarity. No language model required anywhere in the core loop.
Input: a keyword list, imported from CSV or pasted, with a target country, language and device per project.
Collection: for each keyword, fetch the top 10 organic result URLs through a commercial SERP API behind a swappable provider module, never by scraping. Store the URL set with its fetch date. Cache by keyword plus locale plus day and never re-fetch within a day.
Cost, made unavoidable before it is spent: clustering N keywords costs N SERP calls, so show the arithmetic — keyword count times the provider rate — and require acknowledgement before collection starts. Enforce a hard budget that pauses rather than overruns. A user who imports 5,000 keywords without seeing the bill first has been failed by the tool.
Clustering, which is the point of this build:
1. Normalise each result URL — strip the scheme, drop tracking parameters, lowercase the host, remove a trailing slash — so the same page counted twice does not break the comparison.
2. Two keywords are connected when their result sets share at least K URLs, with K configurable and defaulting to 3.
3. Build the graph of those connections and extract clusters. Offer two modes: connected components (loose, larger clusters) and a stricter mode requiring every pair inside a cluster to meet the threshold. Show both and let the user compare — the right threshold is a judgement about how much content consolidation they want, not a fact.
4. Name each cluster by the keyword with the most connections inside it, which is a better label than the highest-volume keyword and needs no volume data.
Why this rather than embeddings, stated in the README: two keywords can read almost identically and return entirely different results, and two that read nothing alike can return the same page. Search-result overlap is a direct measurement of how a search engine treats the pair; semantic similarity is a guess about it.
Inspection: for any cluster, show the shared URLs and how many keywords each one covers. For any pair, show both result sets side by side with the overlap highlighted. A cluster you cannot inspect is a cluster you cannot trust to justify merging two articles.
Outputs: clusters as CSV (keyword, cluster id, cluster name, connection count), a per-cluster brief listing its keywords and shared URLs to research, and a full JSON export.
Re-clustering: change the threshold or the mode and re-cluster from cached SERPs, instantly and at no cost. Threshold experimentation is the main way this tool gets used, so it must be free after collection.
Out of scope: keyword discovery or search volume from any source, content drafting, plagiarism checking, and rank tracking.
Opening prefills the prompt — press enter to run it.
Exhibit B — What you lose
- B.1 the keyword database you browse to build the list in the first place
- B.2 search volume and difficulty for anything in a cluster
- B.3 the plagiarism checking bundled into the workflow
- B.4 the team content workflow with assignments and states
Exhibit C — Why people still pay: workflow, data, and model tuning
Because clustering needs a list, and the list comes from a keyword database nobody can rebuild. WriterZen's price covers discovery as much as the grouping.
Questions
Can I import my WriterZen keyword lists?
Yes, as CSV, and the keywords are all this build needs. Its clusters do not import because they were computed against a different threshold and a different SERP snapshot; re-clustering is a fresh collection run.
Why not cluster with embeddings, which are free?
Because they answer a different question. "Best running shoes" and "top running shoes" are semantically near-identical and may return substantially different results; if so they are two pages, not one. Overlap measures what the search engine does, and that is the decision the clustering exists to inform.
What does it cost to run?
One SERP call per keyword, so clustering a thousand keywords is one to three dollars, once. Re-clustering at different thresholds afterwards is free, which is why the prompt caches aggressively and separates collection from clustering.
What is the one thing that does not survive the rebuild?
The list. Clustering is worthless without keywords to cluster, and the keyword database is the part that requires a crawl of the search results at a scale no personal build reaches. You need a source for the input before this tool has anything to do.
Related tools
Receipt