VibeKoding / Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat / Data Models: A Complete Overview โ€” Document, Graph, Time-Series, and VectorData Models: A Complete Overview โ€” Document, Graph, Time-Series, and Vector
VK

Data Models: A Complete Overview โ€” Document, Graph, Time-Series, and VectorData Models: A Complete Overview โ€” Document, Graph, Time-Series, and Vector

๐Ÿ“š Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat ๐ŸŒ Dual Bahasa (ID / EN) โšก VibeKoding Native

Ensiklopedia VibeKoding: Data Models: A Complete Overview โ€” Document, Graph, Time-Series, and Vector.Ensiklopedia VibeKoding: Data Models: A Complete Overview โ€” Document, Graph, Time-Series, and Vector.

๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

Why can't you just stuff all your data into MySQL tables? When your data is a social network graph, millions of sensor readings per second, or semantic vectors for AI to understand, relational tables fall short. Different data shapes require different modeling approaches.Why can't you just stuff all your data into MySQL tables? When your data is a social network graph, millions of sensor readings per second, or semantic vectors for AI to understand, relational tables fall short. Different data shapes require different modeling approaches.

------

1. Beyond Relational: Motivation for needing Other Data Models1. Beyond Relational: Motivation for needing Other Data Models

Relational databases (MySQL, PostgreSQL) organize data with "tables + rows + columns," suitable for structured, well-defined business data. But real-world data comes in far more forms than just this:Relational databases (MySQL, PostgreSQL) organize data with "tables + rows + columns," suitable for structured, well-defined business data. But real-world data comes in far more forms than just this:

Data ShapeRelational Pain PointBetter Model
User profiles (flexible fields, nested structures)Frequent ALTER TABLE, many NULL columnsDocument Model
Social networks (friends of friends of friends)Multi-level JOIN performance degrades exponentiallyGraph Model
Monitoring metrics (millions of writes per second)Write bottlenecks, historical data bloatTime-Series Model
AI semantic search ("similar meaning" content)Cannot express semantic similarityVector Model
๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

It's not about "replacing" relational databases, but "supplementing" them. Most systems still run their core business on MySQL/PostgreSQL, but introducing specialized data models for specific scenarios can yield orders-of-magnitude performance improvements.It's not about "replacing" relational databases, but "supplementing" them. Most systems still run their core business on MySQL/PostgreSQL, but introducing specialized data models for specific scenarios can yield orders-of-magnitude performance improvements.

------

2. Document Model2. Document Model

2.1 Overview of the Document Model2.1 Overview of the Document Model

The document model stores data as JSON/BSON documents, where each record is a self-contained document that can have different field structures.The document model stores data as JSON/BSON documents, where each record is a self-contained document that can have different field structures.

json
{ "_id": "user_1001", "name": "Zhang San", "tags": ["VIP", "Active"], "address": { "city": "Beijing", "district": "Chaoyang" }, "orders": [ { "id": "o1", "amount": 299 }, { "id": "o2", "amount": 599 } ] }

Key Features:Key Features:

2.2 Document vs. Relational2.2 Document vs. Relational

ComparisonRelational (MySQL)Document (MongoDB)
Data StructureFixed Schema, ALTER TABLE to modifyFlexible Schema, add fields anytime
Nested DataRequires multi-table JOINsEmbedded directly in the document
Cross-record RelationshipsJOINs are powerfulRelationship queries are weaker
Best ForStructurally stable business dataStructurally variable content data

2.3 Typical Use Cases2.3 Typical Use Cases

โš ๏ธ Catatan Keamanan / Peringatanโš ๏ธ Warning / Security Note

"MongoDB doesn't need data structure design" โ€” Wrong! The document model also requires careful design: nesting levels shouldn't be too deep, and frequently updated sub-documents should be split into separate collections."MongoDB doesn't need data structure design" โ€” Wrong! The document model also requires careful design: nesting levels shouldn't be too deep, and frequently updated sub-documents should be split into separate collections.

------

3. Graph Model3. Graph Model

3.1 Overview of the Graph Model3.1 Overview of the Graph Model

The graph model uses Nodes and Edges to represent entities and their relationships. Each node is an entity, each edge is a relationship, and both nodes and edges can carry properties.The graph model uses Nodes and Edges to represent entities and their relationships. Each node is an entity, each edge is a relationship, and both nodes and edges can carry properties.

CODE
(Zhang San) --[follows]--> (Li Si) --[follows]--> (Wang Wu) | | +--------[purchased]----> (iPhone) <--[purchased]--+

3.2 The Graph Model's Killer Feature: Multi-hop Queries3.2 The Graph Model's Killer Feature: Multi-hop Queries

Scenario: Finding "friends of friends of friends" in a social networkScenario: Finding "friends of friends of friends" in a social network

Relational approach (3-level JOIN):Relational approach (3-level JOIN):

sql
SELECT DISTINCT f3.name FROM friends f1 JOIN friends f2 ON f1.friend_id = f2.user_id JOIN friends f3 ON f2.friend_id = f3.user_id WHERE f1.user_id = 1001;

Graph database approach (Cypher query language):Graph database approach (Cypher query language):

cypher
MATCH (me)-[:FOLLOWS*1..3]->(target) WHERE me.name = 'Zhang San' RETURN DISTINCT target.name

Each additional hop in the relational approach adds another JOIN, causing exponential performance degradation. Graph databases traverse relationships via pointers directly, so multi-hop query performance remains nearly unchanged.Each additional hop in the relational approach adds another JOIN, causing exponential performance degradation. Graph databases traverse relationships via pointers directly, so multi-hop query performance remains nearly unchanged.

3.3 Typical Use Cases3.3 Typical Use Cases

------

4. Time-Series Model4. Time-Series Model

4.1 Overview of the Time-Series Model4.1 Overview of the Time-Series Model

The time-series model uses timestamps as the primary axis, specifically optimized for "write in chronological order, query by time range" scenarios.The time-series model uses timestamps as the primary axis, specifically optimized for "write in chronological order, query by time range" scenarios.

CODE
timestamp device cpu_usage memory 2024-01-15 10:00:01 server-01 45% 12.3GB 2024-01-15 10:00:02 server-01 67% 12.5GB 2024-01-15 10:00:03 server-01 92% 14.1GB

4.2 Motivation for Noting Use MySQL for Time-Series Data4.2 Motivation for Noting Use MySQL for Time-Series Data

IssueMySQLTime-Series Database (InfluxDB)
Write SpeedTens of thousands/secMillions/sec
Historical DataManual cleanup, tables keep growingAutomatic expiration policy (TTL)
Aggregation QueriesSlow GROUP BYBuilt-in downsampling (5 sec โ†’ 1 min average)
Storage EfficiencyGeneral-purpose storage, wasted spaceColumnar compression, saving 90% space

4.3 Typical Use Cases4.3 Typical Use Cases

------

5. Vector Model5. Vector Model

5.1 Overview of the Vector Model5.1 Overview of the Vector Model

The vector model converts unstructured data like text, images, and audio into high-dimensional numerical vectors through an Embedding model, then measures semantic similarity by calculating the distance between vectors.The vector model converts unstructured data like text, images, and audio into high-dimensional numerical vectors through an Embedding model, then measures semantic similarity by calculating the distance between vectors.

CODE
"delicious Japanese food" โ†’ Embedding โ†’ [0.82, 0.15, 0.91, 0.33, ...] โ†“ Cosine similarity "Ginza sushi master" โ†’ [0.80, 0.18, 0.89, ...] โ†’ 96% similar "Italian pizza" โ†’ [0.12, 0.85, 0.20, ...] โ†’ 31% similar

5.2 Vector Search vs. Keyword Search5.2 Vector Search vs. Keyword Search

ComparisonKeyword Search (LIKE / Full-text Index)Vector Search
Search MethodExact string matchingSemantic similarity matching
"delicious Japanese food"Can only match text containing "Japanese food"Can find "sushi," "sashimi," "izakaya"
MultilingualNeeds separate handlingCross-language semantic understanding
MultimodalText onlyUnified retrieval across text, images, and audio

5.3 Typical Use Cases5.3 Typical Use Cases

๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

- Standalone Vector Databases: Pinecone, Milvus, Weaviate โ€” focused on vector retrieval, best performance - Traditional Database Extensions: pgvector (PostgreSQL), Atlas Vector Search (MongoDB) โ€” reduce architectural complexity - In-Memory Vector Libraries: FAISS, Annoy โ€” suitable for small-scale, low-latency scenarios- Standalone Vector Databases: Pinecone, Milvus, Weaviate โ€” focused on vector retrieval, best performance - Traditional Database Extensions: pgvector (PostgreSQL), Atlas Vector Search (MongoDB) โ€” reduce architectural complexity - In-Memory Vector Libraries: FAISS, Annoy โ€” suitable for small-scale, low-latency scenarios

------

6. Selection Guide: Approach to choosing a Data Model6. Selection Guide: Approach to choosing a Data Model

What Does Your Data Look Like?Recommended ModelRepresentative Products
Fixed structure, clear relationships (orders, users)RelationalMySQL, PostgreSQL
Flexible structure, deep nesting (content, configs)DocumentMongoDB, DynamoDB
Complex relationships between entities, need multi-hop traversalGraphNeo4j, Amazon Neptune
Write in chronological order, query by time rangeTime-SeriesInfluxDB, TimescaleDB
Unstructured data, need semantic similarity searchVectorPinecone, Milvus, pgvector
๐Ÿ“– Konsep Penting๐Ÿ“– Core Concept

Modern systems typically use multiple models together: - Core business on PostgreSQL (relational) - User behavior logs on InfluxDB (time-series) - AI knowledge base on Milvus + pgvector (vector) - Recommendation engine on Neo4j (graph) Don't try to find "one database to solve all problems" โ€” instead, let each type of data find its most suitable home.Modern systems typically use multiple models together: - Core business on PostgreSQL (relational) - User behavior logs on InfluxDB (time-series) - AI knowledge base on Milvus + pgvector (vector) - Recommendation engine on Neo4j (graph) Don't try to find "one database to solve all problems" โ€” instead, let each type of data find its most suitable home.