Recap from last week
I’ve just about finished building social features in the reader. That means at the end of every chapter you read, you can answer questions to challenge your understanding of what you read - just like my English teacher used to try to get me to do. My hypothesis is that this is the best way to get people’s thoughts onto a screen (of course, there are other thoughts like sharing highlights, taking a photo of your book etc).
When you answer questions, you unlock other peoples’ answers and can discuss with them.
This is somewhat technically challenging, but I’ve tried to make it as easy as possible by making all answers public (for now).
I’ve also been working on scaling questions across a lot of books (next book has 95 chapters). The first book questions were written with philosophy professors, and I need to work out if AI can get close to this level.
Tools and tricks
The difference in session to session output is quite remarkable: Claude Code can go from incredibly helpful to useless with a simple restart, even when I load up all the same files to start the context. I’m not sure why this is but trying to work it out.
When you get a good session, I’m finding it helpful to ask Claude to review the relevant context files I uploaded and if it can improve them, please do so.
When they get too big, you can also ask it to optimise them.
Gemini vs Claude vs ChatGPT: Because I’ve started generating some end of chapter questions with AI, it’s been a good test of which model performs the best.
I create .md files with examples of good questions, context about the project, context about what sort of questions should be generated and .html files of the chapter text and ask the AI to write end of chapter questions. This way every model has the exact same context.
The responses are remarkably similar. Almost all models will pick out the same key passages of the text and ask a question about it. Most models provide 1 or 2 good questions out of 6 per chapter.
Claude Code was the best but this had a bit more context given all my work is done in it.
Going to try Codex this week.
It shows though that the more (and better) context you can give the model, it really does improve.
Reading
I found this article to be very helpful in building evals (ways to check AI is working well): https://www.braintrust.dev/docs/best-practices/scorers
A quick reality check as well of how quick AI is moving (it’s easy to forget that 3 years ago we were amazed it could write a 4 sentence poem): https://www.oneusefulthing.org/p/three-years-from-gpt-3-to-gemini
Next up
Getting the MVP ready for use.
Annoying things like building the landing page, an onboarding flow and various automated emails.
Testing it all thoroughly. Hopefully by the end of the week I’ll be testing the live site and all going well, getting it ready to promote.