Modeling Data with Prisma and MongoDB: Embedding vs Referencing
How to design MongoDB collections when you use Prisma — when to embed, when to reference, how relations work, and the constraints that come with Prisma's MongoDB connector.
MongoDB has no schema enforcement by default and no joins in the relational sense, which makes it easy to start and easy to design badly. Prisma adds a typed schema and a relation layer on top. Using them together works well, provided you understand which of Prisma's relational ideas map cleanly onto documents and which do not.
Setting up the datasource
Prisma's MongoDB connector uses the same schema language as the SQL connectors, with a few differences. Every model needs an @id mapped to MongoDB's _id, using ObjectId:
datasource db {
provider = "mongodb"
url = env("DATABASE_URL")
}
generator client {
provider = "prisma-client-js"
}
model User {
id String @id @default(auto()) @map("_id") @db.ObjectId
email String @unique
name String
}There are no SQL-style migrations. You apply schema changes with prisma db push, which creates collections and indexes. That is convenient, and it also means you own the discipline around changing shapes in a live database.
The core question: embed or reference?
In a relational database you normalise by default. In MongoDB the question is real, and the answer depends on how the data is read and how it grows.
Embed when the child data:
- is always read together with the parent,
- belongs to exactly one parent,
- is bounded in size,
- does not need to be queried independently.
Reference when the child data:
- is large or grows without bound,
- is shared between parents,
- is queried or updated on its own,
- has its own lifecycle and permissions.
A quiz question's answer options are a good embedding case: they are meaningless without the question, they are few, and you always load them together. A student's attempts at that quiz are a good referencing case: they grow forever, are queried by student, and have their own timestamps.
Embedding with composite types
Prisma supports embedded documents through type:
type Option {
label String
isCorrect Boolean
}
model Question {
id String @id @default(auto()) @map("_id") @db.ObjectId
prompt String
options Option[]
}Reads return the options inline, with no second query. Updates to a single option mean rewriting part of the document, so embedding suits data that changes together.
Referencing with relations
References use an ObjectId field and a relation. Both sides are declared; only the foreign key is stored:
model Test {
id String @id @default(auto()) @map("_id") @db.ObjectId
title String
sections Section[]
}
model Section {
id String @id @default(auto()) @map("_id") @db.ObjectId
title String
order Int
testId String @db.ObjectId
test Test @relation(fields: [testId], references: [id])
@@index([testId, order])
}Prisma resolves include: { sections: true } with an additional query behind the scenes. That is fine, but it is not a SQL join — which brings us to the constraints.
Constraints worth knowing
- Referential integrity is not enforced by the database. MongoDB has no foreign keys. Prisma emulates relation behaviour in the client; deleting a parent does not automatically make the database remove orphaned children unless you specify referential actions and use them through Prisma. Do deletes through one code path.
- Index what you query. Relation fields are not indexed for you. Add
@@indexfor foreign keys and for the sort order you use on lists. - Unique constraints work, including compound ones:
@@unique([userId, testId, attemptNumber]). - Transactions need a replica set. Multi-document transactions require MongoDB to run as a replica set (which Atlas does by default). Local single-node development setups often do not, so test transactional code against a replica-set configuration.
- Documents are capped at 16 MB. Unbounded embedded arrays are a time bomb. If an array can grow with usage, reference instead.
Designing around the read path
Start from the screens you need to render, not from the entities. For a test-taking screen, list what a single request must return: the test, its sections, the passages and the questions. If that is read together every time and rarely edited, you may embed passages in sections. If different teams edit questions independently and reuse them across tests, reference them.
A useful exercise is to write the queries first:
const test = await prisma.test.findUnique({
where: { id },
include: {
sections: {
orderBy: { order: "asc" },
include: { questions: { orderBy: { order: "asc" } } },
},
},
});If a screen needs five nested includes, either the data wants to be embedded or you need a purpose-built read model.
Recording attempts and results
Anything user-generated and append-only belongs in its own collection with indexes on the access path. A common shape for exam-style data:
model Attempt {
id String @id @default(auto()) @map("_id") @db.ObjectId
userId String @db.ObjectId
testId String @db.ObjectId
startedAt DateTime @default(now())
finishedAt DateTime?
score Float?
@@index([userId, startedAt])
@@index([testId])
}Storing derived values such as a final score on the attempt itself is a deliberate denormalisation. It makes dashboards cheap, at the cost of having to recompute when the scoring rules change. Decide which values are source data and which are caches, and write that down.
Changing a schema safely
Without migrations, changes are your responsibility:
- Add fields as optional first. Existing documents do not have them.
- Backfill with a script that is safe to re-run.
- Then make the field required, if you need that.
- Deploy code that tolerates both shapes before deploying code that requires the new one.
Treat renames and type changes as multi-step operations, never a single edit.
A short summary
- Embed what is read together, bounded and owned by one parent.
- Reference what grows, is shared or has its own lifecycle.
- Index relation fields and list-ordering fields yourself.
- Remember there are no database-enforced foreign keys.
- Evolve schemas in additive steps.
For how these rules play out in a larger domain, Designing an LMS Data Model applies them to courses, lessons, quizzes and progress. The IELTS Mock and E2A Learning projects both use this stack.
Related articles
- Designing an LMS Data Model: Courses, Lessons, Quizzes and Progress
A practical data model for a learning platform — content hierarchy, free previews, quiz attempts, past-paper practice and progress tracking — with the trade-offs behind each choice.
- Structuring a Next.js App Router Project for Production
A practical layout for Next.js App Router projects — routes, data access, shared components and configuration — and the reasoning behind each boundary.