Nice constraint. I've been building in this space for a while now and the lossiness question is a real one. The accessibility tree works for a lot of things but the more I worked on it the more stuff I found needed pixels.
Hmm. What would be really helpful for me would be an enhancement of the TypeWhisper app for dictation that would be able to get the context of what I am dictating into and send it along with the prompt to an LLM.
Probably much simpler and much less of a privacy problem (I run my own local LLM for that purpose so that nothing leaves the machine).
So the markdown is scaffolding the app generates deterministically.
The ## heading is built from the block's timestamps, the app name and the window title. The frontmatter is per-day boilerplate. The file:/url: line is the window's backing document where the app exposes one. The captured text underneath is written exactly as the tree handed it over: plain lines, no reconstruction.
That lossiness is also why the file/URL references exist. Trying to rebuild a document's formatting from its accessibility tree is a losing game, so instead each block records where the real document lives, and the LLM reading the file can open the original if the fragments aren't enough. "Plain markdown" was meant as "a markdown file you can open anywhere", not "faithful markdown conversion of what you saw"
A cool project, but feels like fundamentally the wrong approach when what you do most likely leaves a string of structured digital footprints anyway. I’ve got an agent that fills out my timesheets by looking at git commits, agent history, Slack messages, emails, and time-tracker tickets. I guess I could add relevant web-browsing?
When I built rem, I spent significant effort getting screenshot -> ocr + screenshot -> ffmpeg loop energy efficient, but it definitely is more expensive than accessibility API.
You also save a lot of disk space and writes to disk.
That being said, you lose the cool swipe to go back in time and search through history and visually see, features.
And situations where accessibility isn't supported.
And as others have mentioned, built in ocr is definitely better than tesseract.
Yes, I know. Like I said: “Project reasons aside”. I’m not suggesting OCR for this, I’m imparting the general information to be used in other situations that in macOS you can OCR without requiring third-party tools.
I experimented with this exact same approach earlier this year.
It's barely sufficient, because, bluntly, most apps just aren't wired up right.
So you end up having to hand code a lot of specific profiles for specific apps to make this work well, and even then, you don't quite get the right level of detail to make it work out.
Will try this app, to see if it improved on my own approach, but man, the hope levels are low.
I hope in near future _that_ layer of abstraction -- looking at a fairly standard application window with minor UI variations and reasoning about what area / labels within the UI mean what (possibly paired with app documentation) -- could probably become a light-weight fine-tuned vision model it itself that can run fully locally.
I also saw that HeyClicky started doing something similar but end up removing from the product.
This is pretty magical tbh
Also added MIT license and added some more pruning to reduce initial capture by around 10%.
YMMV.
Probably much simpler and much less of a privacy problem (I run my own local LLM for that purpose so that nothing leaves the machine).
> It writes plain markdown
Where are the formatting decisions coming from?
The ## heading is built from the block's timestamps, the app name and the window title. The frontmatter is per-day boilerplate. The file:/url: line is the window's backing document where the app exposes one. The captured text underneath is written exactly as the tree handed it over: plain lines, no reconstruction.
That lossiness is also why the file/URL references exist. Trying to rebuild a document's formatting from its accessibility tree is a losing game, so instead each block records where the real document lives, and the LLM reading the file can open the original if the fragments aren't enough. "Plain markdown" was meant as "a markdown file you can open anywhere", not "faithful markdown conversion of what you saw"
You also save a lot of disk space and writes to disk.
That being said, you lose the cool swipe to go back in time and search through history and visually see, features.
And situations where accessibility isn't supported.
And as others have mentioned, built in ocr is definitely better than tesseract.
It's barely sufficient, because, bluntly, most apps just aren't wired up right.
So you end up having to hand code a lot of specific profiles for specific apps to make this work well, and even then, you don't quite get the right level of detail to make it work out.
Will try this app, to see if it improved on my own approach, but man, the hope levels are low.
UPDATE: As i just checked it is also using Accessibility API. So i guess we will have access to the same set to data.