Show HN: LatticeDB – Like SQLite but for graph databases
We have been using graph DBs more and more at work. I found them painful to work with locally and decided to try and build something better.
128 points by smiths1999 - 37 commentsWe have been using graph DBs more and more at work. I found them painful to work with locally and decided to try and build something better.
128 points by smiths1999 - 37 comments
Given that on-disk data structures are similar to SQLite, I expect the competition from other "graph on sqlite" projects when they co-opt the techniques in LatticeDB.
I'm currently building a personal knowledge graph server a mix of Notion's custom entities via JSON schema and Obsidian markdown+backlinked references. It's working well, but I suspect your product might be a better fit.
I do have one question regarding permissions: how would you recommend modeling a hierarchical access system in a graph database? Specifically, if a user is granted access to a document, they should automatically have access to all its child documents within that workspace. Is there a standard way to model this 'subtree' permission logic, or perhaps a more efficient approach you'd suggest?
Really impressed with the product good luck with it!
Anyways, thanks for checking it out! Really appreciate it. Good luck with your project!
My team and I built it over the years, and it's open source (AGPL). Here is how it works: https://community.qbix.com/t/qbix-streams-as-a-graph-databas...
Our database abstraction layer (and optional ORM) has been battle tested in production for millions of users, and has 3 adapters: Sqlite, Postgres, MySQL/MariaDB. And recently it even added vector search for ranking results by similarity: https://github.com/Qbix/Platform/tree/main/platform/classes/...
Documentation for the database layer is here: https://qbix.com/platform/guide/database
PS: If you do use a relational database for storing graph data, you're going to have a lot of duplication in some public keys. I highly recommend putting ZFS underneath, to help with deduplication. ZFS uses zstd, developed at facebook, and also can encrypt your data at rest (don't use the relational database to do the encryption, otherwise deduplication doesn't work).
https://huggingface.co/datasets/ladybugdb/wikidata-20260401
In general though, the goal for this was single writer multiple readers. That was a design decision to keep things simple.
As for the tool, it scratches an itch I've been having, I'll give it a go soon.
The initial phases of building I would build out piece by piece. For example, building out the file system interactions I would have claude build a feature and explain how it worked in an educational manner (e.g., like it was a section in a book on latticedb internals). I would then read through the code. This was a great way to learn and build, simultaneously.
In the later stages, where the features and work was more complex, I would spend more time discussing, instructing, and verifying, but less time understanding the actual implementation. I'll give you an example. It's been a long time since I've handwritten SIMD code. I could try and review claudes output, but I am certain I'd miss any subtle bugs that may exist. I found it more productive to assume the code was right and focus on thinking about how I would verify that. Benchmarking, playing with latticedb, etc. were my primary tools for verifying the work. I could run a benchmark and see performance was great. Then I'd explore the test vectors and realize they were trivial, completely invalidating the benchmark results. So we would go back to the drawing board, create a new benchmark set, see results weren't great, and evaluate what was wrong with the implementation. Sometimes features would take days to get out just because of the iteration loop.
https://duckdb.org/community_extensions/extensions/duckpgq
One of the motivating use cases for me was experimenting with agentic memory. I use latticedb as the backing data store. Finding related memories is traversing the graph (kind of like graph RAG).
What are some of the scales of the data you've been able to test this design on so far?
What was the most interesting part of designing it for you?
Most interesting part is a tough one. From a learning perspective the beginning was incredibly interesting because I was spending a lot of time learning about how other DBs work. Even something as relatively simple as writing to disk had a lot more complexity to it than I initially anticipated.
I used LLMs extensively in building this, and the other interesting part was seeing how they failed. I've always been a proponent that tests are no guarantee of quality code, and working with LLMs has only reinforced it. They often write superficial tests. Sometimes a suite of tests would pass, but when I would actually play around with the feature it was clearly broken. LLMs certainly enabled me to build something of this scope, but it was far from "build a graph DB and notify me when you are done"
However, I really miss the content posted by the KuzuDB team on their YouTube channel.
Couple of corrections:
* LadybugDB has revamped the Kuzu WAL design. It shouldn't be hard to build WAL based replication
* 19ms vs 39us - like the author says these are vastly different systems and the benchmark methodology may not be comparable.
We've mostly focused on query plan optimizations, not so much the micro query operator optimizations.
The 0.20.x end of the month release should have some interesting optimizations.
For those into the `pip install ...` flow and kuzu, is gfql: we started around the same time in a non-VC-funded oss manner with overlap in key architectural ideas:
- cpu columnar vectorized engine + optionally the only open source gpu engine mode for bigger graphs / faster queries
- removes the need for a database / file: pure compute-tier engine you can write to parquet/json if you want, plays with parallel reader/writers in simple ways b/c that, and TBD iceberg
- adds full graph analytic pipeline support, eg, for feature engineering in real-time fraud & memory pipelines
- also millisecond/submillisecond times on small graphs like that small 100K edge graph benchmark
Main box not formally checked is streaming. Funny enough, we're designed for GPU firehose workloads, so would be fun to demo and see what gaps are left.