๐ Milvus v2.4.3์ Metadata Filtering ์๋ก์ด ๊ธฐ๋ฅ
Milvus v2.4.3์์๋ ์ ์ฒด ๋ฌธ์์ด ๋ฉํ๋ฐ์ดํฐ ๋งค์นญ์ ๋์ ํ์ต๋๋ค! ๐ ์ด์ ์ ๋์ฌ, ์ค์, ์ ๋ฏธ์ฌ ๋๋ ๋ฌธ์ ์์ผ๋์นด๋ ๊ฒ์์ ์ฌ์ฉํด ๋ฌธ์์ด์ ๋งค์นญํ ์ ์์ต๋๋ค.
# Prefix example, matches any string starting with โTheโ.
expression='title like "The%"'
# Infix example, matches any string with the word โtheโ anywhere in the sentence.
expression='title like "%the%"'
# Postfix example, matches any string ending with โRyeโ.
expression='title like "%Rye"'
# Single character wildcard example, matches any one single character at a specific position.
expression='title like "Flip_ed"'
์ด์ ๋ธ๋ก๊ทธ์์๋ ์ ๋์ฌ ๋ฌธ์์ด ๋งค์นญ์ ๋ํด์๋ง ์ด์ผ๊ธฐํ์ต๋๋ค. ํ์ง๋ง Milvus v2.4.3๋ถํฐ๋ ๋ชจ๋ ๋ณํ์ด ๊ฐ๋ฅํ๋ฉฐ, ๋ฐฐ์ด ๊ฐ๋ ์ ํํ ์ผ์น์ํค๊ฑฐ๋ ๋ฐฐ์ด์ ์ด๋ค ์์๋ ์ผ์นํ๋์ง ํ์ธํ๋ ๋ฐฉ์(contains_any())์ผ๋ก ์ฌ์ฉํ ์ ์์ต๋๋ค. ๐๏ธ๐
์ด๋ฌํ ์ ๋ฐ์ดํธ๋ก ๋ฉํ๋ฐ์ดํฐ ํํฐ๋ง์ด ๋์ฑ ๋ค์ฌ๋ค๋ฅํ๊ณ ๊ฐ๋ ฅํด์ก์ต๋๋ค!
์์๋ฅผ ํตํด ์ด๋ฅผ ๋ ๋ช ํํ ํด๋ณด๊ฒ ์ต๋๋ค. ์ด ๋ธ๋ก๊ทธ์์๋ Kaggle์์ ๋ค์ด๋ก๋ํ IMDB ์ํ ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ๊ฒ ์ต๋๋ค.
# Import common libraries.
import sys, os, time, pprint
import pandas as pd
# Read CSV data.
df = pd.read_csv("data/original_data.csv")
# Shortcut the data for demo.
df = df.tail(200)
display(df.head())
๊ฐ ์ํ์๋ ์ค๋ช ๊ณผ ๋ฆฌ๋ทฐ๊ฐ ํฌํจ๋ โtextโ ํ๋๊ฐ ์์ต๋๋ค. ๐
๊ฐ ์ํ์๋ ๊ฐ๋ด ์ฐ๋, ํ์ , ์ฅ๋ฅด, ๋ฐฐ์ฐ, ํค์๋ ๋ชฉ๋ก๊ณผ ๊ฐ์ ๋ฉํ๋ฐ์ดํฐ๊ฐ ํฌํจ๋์ด ์์ต๋๋ค. ๐ฌโญ๏ธ๐
๊ฐ โํโ์ ์ํ ๋ฆฌ๋ทฐ ํ ์คํธ ์ฒญํฌ, ๊ทธ ๋ฒกํฐ ํํ, ๊ทธ๋ฆฌ๊ณ movie_id, ์ํ ์ ๋ชฉ, ํฌ์คํฐ ๋งํฌ, ์ฅ๋ฅด, ๋ฐฐ์ฐ์ ๊ฐ์ ๋ฉํ๋ฐ์ดํฐ๋ฅผ ๋ํ๋ ๋๋ค.
์ผ๋ฐ์ ์ธ RAG ํจํด์ ๋ฐ๋ฅด๋ฉด: ๐๐ฌ
Milvus์ ์ฐ๊ฒฐ: ๋จผ์ Milvus์ ๋ก์ปฌ ๋ฐฐํฌํ์ธ Milvus Lite์ ์ฐ๊ฒฐํฉ๋๋ค. ์ด๋ ๋ฒกํฐ๋ฅผ ์ ์ฅํ๊ณ ๊ด๋ฆฌํ๊ธฐ ์ํ ๋ฐ์ดํฐ๋ฒ ์ด์ค์ ๋๋ค. ๐ฅ๏ธ๐
์ํ ํ ์คํธ๋ฅผ ๋ฒกํฐ๋ก ๋ณํ: ์ค๋ช ๊ณผ ๋ฆฌ๋ทฐ๊ฐ ํฌํจ๋ ๊ฐ ์ํ์ ํ ์คํธ ํ๋๋ฅผ ๊ฐ์ ธ์ ๋ฒกํฐ๋ก ๋ณํํฉ๋๋ค. ์ด๋ฅผ ์ํด HuggingFace ๋ชจ๋ธ BAAI/bge-large-en-v1.5์ ์ฌ์ฉํฉ๋๋ค. ๐ง โก๏ธ๐
๐ฅ๐ ๋ฒกํฐ์ ๋ฉํ๋ฐ์ดํฐ๋ฅผ Milvus์ ์ฝ์ : ์ด ๋ฒกํฐ๋ฅผ ์๋ณธ ํ ์คํธ(โ์ฒญํฌโ๋ผ๊ณ ํจ) ๋ฐ ํด๋น ๋ฉํ๋ฐ์ดํฐ(์ฐ๋, ํ์ , ์ฅ๋ฅด ๋ฑ)์ ํจ๊ป Milvus์ ์ฝ์ ํฉ๋๋ค. ๐ฅ๐
์ฌ์ฉ์ ์ฟผ๋ฆฌ ์ฒ๋ฆฌ: ๋์ผํ ์๋ฒ ๋ฉ ๋ชจ๋ธ์ ์ฌ์ฉํด ์ฌ์ฉ์์ ์ฟผ๋ฆฌ๋ฅผ ๋ฒกํฐ๋ก ๋ณํํฉ๋๋ค. ๊ทธ๋ฐ ๋ค์ Approximate Nearest Neighbors ๊ฒ์์ ์คํํ์ฌ ์ฟผ๋ฆฌ ๋ฒกํฐ์ ๊ฐ์ฅ ๊ฐ๊น์ด ๋ฐ์ดํฐ ๋ฒกํฐ๋ฅผ ์ฐพ์ต๋๋ค. ๐๐ฌ
์ ์ฒด ์ฝ๋๋ ์ GitHub์์ ํ์ธํ ์ ์์ต๋๋ค.
๋จผ์ Milvus์ ์ฐ๊ฒฐํฉ๋๋ค. Pymilvus๋ฅผ pip๋ก ์ค์นํด์ผ ํฉ๋๋ค. (๋ก์ปฌ ํ์ผ ์ด๋ฆ๋ง ์ง์ ํ๋ฉด ๋ก์ปฌ ๋ฒกํฐ ๋ฐ์ดํฐ๋ฒ ์ด์ค์ธ Milvus Lite๋ฅผ ์ฌ์ฉํฉ๋๋ค. Docker๋ K8s๋ก ๋ฐฐํฌ๋ ๋ค๋ฅธ Milvus ๋๋ ์์ ๊ด๋ฆฌํ Zilliz Cloud๊ฐ ์๋ค๋ฉด URI์ Token์ ์ง์ ํด ์ฐ๊ฒฐํ ์ ์์ต๋๋ค. ๋๋จธ์ง ์ฝ๋๋ ๋์ผํ๊ฒ ์๋ํฉ๋๋ค.)
# !python -m pip install -U pymilvus
import pymilvus
# Connect a client to the Milvus Lite server.
from pymilvus import MilvusClient
client = MilvusClient("milvus_demo.db")
๋ค์์ผ๋ก, ์ํ ๋ฆฌ๋ทฐ๊ฐ ํฌํจ๋ ํ ์คํธ ์ด์ ์ฒญํฌ๋ก ๋๋๊ณ ์๋ฒ ๋ฉํ์ฌ ๋ฒกํฐ๋ก ๋ณํํฉ๋๋ค. ์ด๋ฅผ ์ํํ๋ ๋ฐฉ๋ฒ์ ๋ํ ์์๋ ๋ง์ ๋ฆฌ์์ค์์ ์ ๊ณตํ๋ฏ๋ก, ์ฌ๊ธฐ์๋ ์ฝ๋๋ฅผ ๋ค์ ๋ณด์ฌ๋๋ฆฌ์ง ์๊ฒ ์ต๋๋ค. ์๋์์๋ ์ฒญํฌํ๋ ํ ์คํธ, ๋ฒกํฐ ํํ, ๋ฉํ๋ฐ์ดํฐ๋ฅผ ์กฐํฉํ๊ณ ๋ฐ์ดํฐ๋ฅผ Milvus์ ์ฝ์ ํ๋ ๋ฐฉ๋ฒ์ ๋ณด์ฌ๋๋ฆฝ๋๋ค.
# Create chunk_list and dict_list in a single loop
dict_list = []
for id, title, chunk, vector, poster_url, director,\
genres, actors, keywords, film_year, rating in zip(
df.id, df.Name, chunks, converted_values, df.PosterLink,
df.Director, df.Genres, df.Actors, df.Keywords,
df.MovieYear, df.RatingValue):
# Assemble embedding vector, original text chunk, metadata.
chunk_dict = {
'movie_index': id,
'title': title,
'chunk': chunk.page_content,
'poster_url': poster_url,
'director': director,
'genres': genres,
'actors': actors,
'keywords': keywords,
'film_year': film_year,
'rating': rating,
'vector': vector,
}
dict_list.append(chunk_dict)
# Insert data into the Milvus collection.
print("Start inserting entities")
start_time = time.time()
client.insert(
COLLECTION_NAME,
data=dict_list,
progress_bar=True)
end_time = time.time()
print(f"Milvus insert time for {len(dict_list)} vectors: ", end="")
print(f"{np.round(end_time - start_time, 2)} seconds")
์ด์ ๋ฐ์ดํฐ๊ฐ Milvus์ ์์ผ๋ฏ๋ก ๊ฒ์ํ ์ค๋น๊ฐ ๋์์ต๋๋ค!
๋ฌธ์์ด ๋ฉํ๋ฐ์ดํฐ ํํฐ๋ฅผ ์ฌ์ฉํ ๊ฒ์
๋ก๋ด์ด ๋ฑ์ฅํ๋ ๋์คํ ํผ์์ ๋ฏธ๋๋ฅผ ๋ค๋ฃฌ ๋์ ํ์ ์ SF ์ํ๋ฅผ ์ฐพ๊ณ ์ถ๋ค๊ณ ๊ฐ์ ํด ๋ด ์๋ค. ์์ ๋ฐ์ดํฐ์๋ ์ด ๊ฒ์์ ์ฌ์ฉํ ์ ์๋ ๋ฉํ๋ฐ์ดํฐ๊ฐ ์์ต๋๋ค.
๋ค์์ ํผ์ง ๋ฌธ์์ด ๋งค์น๋ฅผ ์ฌ์ฉํ ๋ฉํ๋ฐ์ดํฐ ํํฐ๋ง์ ์์ ๋๋ค. ์๋์์๋ ๊ฐ ๊ฒ์ ํ ๋ฉํ๋ฐ์ดํฐ๋ฅผ ๋ ์ฝ๊ฒ ํ์ํ๊ธฐ ์ํด ๊ธฐ๋ณธ Milvus search API๋ง ๋ํํ์ต๋๋ค.
SAMPLE_QUESTION = "Dystopia science fiction with a robot."
TOP_K = 1
# Metadata filters.
expression='rating >= 7'
# Infix string match.
expression=expression + ' && title like "%Panther%"'
formatted_results, contexts, context_metadata = \
mc_run_search(SAMPLE_QUESTION, expression, TOP_K)
๋ฆฌ์์ค ๋ฐ ์ถ๊ฐ ์๋ฃ
Milvus ๋น ๋ฅธ ์์ ๊ฐ์ด๋
๋ฐฐ์ด ํ๋ ์ฌ์ฉ | Milvus ๋ฌธ์
https://github.com/milvus-io/pymilvus/blob/2.4/examples/fuzzy_match.py
https://milvus.io/docs/boolean.md#Usage
https://milvus.io/docs/single-vector-search.md#Filtered-search
๊ณ์ ์ฝ๊ธฐ

Why We Built Vector Lakebase: Rethinking Unstructured Data Architecture for AI
Vector Lakebase: a unified, lake-native data foundation for AI workloads โ and an answer to what happens after vector databases succeed.

Vector Databases vs. Document Databases
Use a vector database for similarity search and AI-powered applications; use a document database for flexible schema and JSON-like data storage.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Explore DeepSeek-VL2, the open-source MoE vision-language model. Discover its architecture, efficient training pipeline, and top-tier performance.



