๐ Milvus v2.4.3์ Metadata Filtering ์๋ก์ด ๊ธฐ๋ฅ
Milvus v2.4.3์์๋ ์ ์ฒด ๋ฌธ์์ด ๋ฉํ๋ฐ์ดํฐ ๋งค์นญ์ ๋์ ํ์ต๋๋ค! ๐ ์ด์ ์ ๋์ฌ, ์ค์, ์ ๋ฏธ์ฌ ๋๋ ๋ฌธ์ ์์ผ๋์นด๋ ๊ฒ์์ ์ฌ์ฉํด ๋ฌธ์์ด์ ๋งค์นญํ ์ ์์ต๋๋ค.
# Prefix example, matches any string starting with โTheโ.
expression='title like "The%"'
# Infix example, matches any string with the word โtheโ anywhere in the sentence.
expression='title like "%the%"'
# Postfix example, matches any string ending with โRyeโ.
expression='title like "%Rye"'
# Single character wildcard example, matches any one single character at a specific position.
expression='title like "Flip_ed"'
์ด์ ๋ธ๋ก๊ทธ์์๋ ์ ๋์ฌ ๋ฌธ์์ด ๋งค์นญ์ ๋ํด์๋ง ์ด์ผ๊ธฐํ์ต๋๋ค. ํ์ง๋ง Milvus v2.4.3๋ถํฐ๋ ๋ชจ๋ ๋ณํ์ด ๊ฐ๋ฅํ๋ฉฐ, ๋ฐฐ์ด ๊ฐ๋ ์ ํํ ์ผ์น์ํค๊ฑฐ๋ ๋ฐฐ์ด์ ์ด๋ค ์์๋ ์ผ์นํ๋์ง ํ์ธํ๋ ๋ฐฉ์(contains_any())์ผ๋ก ์ฌ์ฉํ ์ ์์ต๋๋ค. ๐๏ธ๐
์ด๋ฌํ ์ ๋ฐ์ดํธ๋ก ๋ฉํ๋ฐ์ดํฐ ํํฐ๋ง์ด ๋์ฑ ๋ค์ฌ๋ค๋ฅํ๊ณ ๊ฐ๋ ฅํด์ก์ต๋๋ค!
์์๋ฅผ ํตํด ์ด๋ฅผ ๋ ๋ช ํํ ํด๋ณด๊ฒ ์ต๋๋ค. ์ด ๋ธ๋ก๊ทธ์์๋ Kaggle์์ ๋ค์ด๋ก๋ํ IMDB ์ํ ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ๊ฒ ์ต๋๋ค.
# Import common libraries.
import sys, os, time, pprint
import pandas as pd
# Read CSV data.
df = pd.read_csv("data/original_data.csv")
# Shortcut the data for demo.
df = df.tail(200)
display(df.head())
๊ฐ ์ํ์๋ ์ค๋ช ๊ณผ ๋ฆฌ๋ทฐ๊ฐ ํฌํจ๋ โtextโ ํ๋๊ฐ ์์ต๋๋ค. ๐
๊ฐ ์ํ์๋ ๊ฐ๋ด ์ฐ๋, ํ์ , ์ฅ๋ฅด, ๋ฐฐ์ฐ, ํค์๋ ๋ชฉ๋ก๊ณผ ๊ฐ์ ๋ฉํ๋ฐ์ดํฐ๊ฐ ํฌํจ๋์ด ์์ต๋๋ค. ๐ฌโญ๏ธ๐
๊ฐ โํโ์ ์ํ ๋ฆฌ๋ทฐ ํ ์คํธ ์ฒญํฌ, ๊ทธ ๋ฒกํฐ ํํ, ๊ทธ๋ฆฌ๊ณ movie_id, ์ํ ์ ๋ชฉ, ํฌ์คํฐ ๋งํฌ, ์ฅ๋ฅด, ๋ฐฐ์ฐ์ ๊ฐ์ ๋ฉํ๋ฐ์ดํฐ๋ฅผ ๋ํ๋ ๋๋ค.
์ผ๋ฐ์ ์ธ RAG ํจํด์ ๋ฐ๋ฅด๋ฉด: ๐๐ฌ
Milvus์ ์ฐ๊ฒฐ: ๋จผ์ Milvus์ ๋ก์ปฌ ๋ฐฐํฌํ์ธ Milvus Lite์ ์ฐ๊ฒฐํฉ๋๋ค. ์ด๋ ๋ฒกํฐ๋ฅผ ์ ์ฅํ๊ณ ๊ด๋ฆฌํ๊ธฐ ์ํ ๋ฐ์ดํฐ๋ฒ ์ด์ค์ ๋๋ค. ๐ฅ๏ธ๐
์ํ ํ ์คํธ๋ฅผ ๋ฒกํฐ๋ก ๋ณํ: ์ค๋ช ๊ณผ ๋ฆฌ๋ทฐ๊ฐ ํฌํจ๋ ๊ฐ ์ํ์ ํ ์คํธ ํ๋๋ฅผ ๊ฐ์ ธ์ ๋ฒกํฐ๋ก ๋ณํํฉ๋๋ค. ์ด๋ฅผ ์ํด HuggingFace ๋ชจ๋ธ BAAI/bge-large-en-v1.5์ ์ฌ์ฉํฉ๋๋ค. ๐ง โก๏ธ๐
๐ฅ๐ ๋ฒกํฐ์ ๋ฉํ๋ฐ์ดํฐ๋ฅผ Milvus์ ์ฝ์ : ์ด ๋ฒกํฐ๋ฅผ ์๋ณธ ํ ์คํธ(โ์ฒญํฌโ๋ผ๊ณ ํจ) ๋ฐ ํด๋น ๋ฉํ๋ฐ์ดํฐ(์ฐ๋, ํ์ , ์ฅ๋ฅด ๋ฑ)์ ํจ๊ป Milvus์ ์ฝ์ ํฉ๋๋ค. ๐ฅ๐
์ฌ์ฉ์ ์ฟผ๋ฆฌ ์ฒ๋ฆฌ: ๋์ผํ ์๋ฒ ๋ฉ ๋ชจ๋ธ์ ์ฌ์ฉํด ์ฌ์ฉ์์ ์ฟผ๋ฆฌ๋ฅผ ๋ฒกํฐ๋ก ๋ณํํฉ๋๋ค. ๊ทธ๋ฐ ๋ค์ Approximate Nearest Neighbors ๊ฒ์์ ์คํํ์ฌ ์ฟผ๋ฆฌ ๋ฒกํฐ์ ๊ฐ์ฅ ๊ฐ๊น์ด ๋ฐ์ดํฐ ๋ฒกํฐ๋ฅผ ์ฐพ์ต๋๋ค. ๐๐ฌ
์ ์ฒด ์ฝ๋๋ ์ GitHub์์ ํ์ธํ ์ ์์ต๋๋ค.
๋จผ์ Milvus์ ์ฐ๊ฒฐํฉ๋๋ค. Pymilvus๋ฅผ pip๋ก ์ค์นํด์ผ ํฉ๋๋ค. (๋ก์ปฌ ํ์ผ ์ด๋ฆ๋ง ์ง์ ํ๋ฉด ๋ก์ปฌ ๋ฒกํฐ ๋ฐ์ดํฐ๋ฒ ์ด์ค์ธ Milvus Lite๋ฅผ ์ฌ์ฉํฉ๋๋ค. Docker๋ K8s๋ก ๋ฐฐํฌ๋ ๋ค๋ฅธ Milvus ๋๋ ์์ ๊ด๋ฆฌํ Zilliz Cloud๊ฐ ์๋ค๋ฉด URI์ Token์ ์ง์ ํด ์ฐ๊ฒฐํ ์ ์์ต๋๋ค. ๋๋จธ์ง ์ฝ๋๋ ๋์ผํ๊ฒ ์๋ํฉ๋๋ค.)
# !python -m pip install -U pymilvus
import pymilvus
# Connect a client to the Milvus Lite server.
from pymilvus import MilvusClient
client = MilvusClient("milvus_demo.db")
๋ค์์ผ๋ก, ์ํ ๋ฆฌ๋ทฐ๊ฐ ํฌํจ๋ ํ ์คํธ ์ด์ ์ฒญํฌ๋ก ๋๋๊ณ ์๋ฒ ๋ฉํ์ฌ ๋ฒกํฐ๋ก ๋ณํํฉ๋๋ค. ์ด๋ฅผ ์ํํ๋ ๋ฐฉ๋ฒ์ ๋ํ ์์๋ ๋ง์ ๋ฆฌ์์ค์์ ์ ๊ณตํ๋ฏ๋ก, ์ฌ๊ธฐ์๋ ์ฝ๋๋ฅผ ๋ค์ ๋ณด์ฌ๋๋ฆฌ์ง ์๊ฒ ์ต๋๋ค. ์๋์์๋ ์ฒญํฌํ๋ ํ ์คํธ, ๋ฒกํฐ ํํ, ๋ฉํ๋ฐ์ดํฐ๋ฅผ ์กฐํฉํ๊ณ ๋ฐ์ดํฐ๋ฅผ Milvus์ ์ฝ์ ํ๋ ๋ฐฉ๋ฒ์ ๋ณด์ฌ๋๋ฆฝ๋๋ค.
# Create chunk_list and dict_list in a single loop
dict_list = []
for id, title, chunk, vector, poster_url, director,\
genres, actors, keywords, film_year, rating in zip(
df.id, df.Name, chunks, converted_values, df.PosterLink,
df.Director, df.Genres, df.Actors, df.Keywords,
df.MovieYear, df.RatingValue):
# Assemble embedding vector, original text chunk, metadata.
chunk_dict = {
'movie_index': id,
'title': title,
'chunk': chunk.page_content,
'poster_url': poster_url,
'director': director,
'genres': genres,
'actors': actors,
'keywords': keywords,
'film_year': film_year,
'rating': rating,
'vector': vector,
}
dict_list.append(chunk_dict)
# Insert data into the Milvus collection.
print("Start inserting entities")
start_time = time.time()
client.insert(
COLLECTION_NAME,
data=dict_list,
progress_bar=True)
end_time = time.time()
print(f"Milvus insert time for {len(dict_list)} vectors: ", end="")
print(f"{np.round(end_time - start_time, 2)} seconds")
์ด์ ๋ฐ์ดํฐ๊ฐ Milvus์ ์์ผ๋ฏ๋ก ๊ฒ์ํ ์ค๋น๊ฐ ๋์์ต๋๋ค!
๋ฌธ์์ด ๋ฉํ๋ฐ์ดํฐ ํํฐ๋ฅผ ์ฌ์ฉํ ๊ฒ์
๋ก๋ด์ด ๋ฑ์ฅํ๋ ๋์คํ ํผ์์ ๋ฏธ๋๋ฅผ ๋ค๋ฃฌ ๋์ ํ์ ์ SF ์ํ๋ฅผ ์ฐพ๊ณ ์ถ๋ค๊ณ ๊ฐ์ ํด ๋ด ์๋ค. ์์ ๋ฐ์ดํฐ์๋ ์ด ๊ฒ์์ ์ฌ์ฉํ ์ ์๋ ๋ฉํ๋ฐ์ดํฐ๊ฐ ์์ต๋๋ค.
๋ค์์ ํผ์ง ๋ฌธ์์ด ๋งค์น๋ฅผ ์ฌ์ฉํ ๋ฉํ๋ฐ์ดํฐ ํํฐ๋ง์ ์์ ๋๋ค. ์๋์์๋ ๊ฐ ๊ฒ์ ํ ๋ฉํ๋ฐ์ดํฐ๋ฅผ ๋ ์ฝ๊ฒ ํ์ํ๊ธฐ ์ํด ๊ธฐ๋ณธ Milvus search API๋ง ๋ํํ์ต๋๋ค.
SAMPLE_QUESTION = "Dystopia science fiction with a robot."
TOP_K = 1
# Metadata filters.
expression='rating >= 7'
# Infix string match.
expression=expression + ' && title like "%Panther%"'
formatted_results, contexts, context_metadata = \
mc_run_search(SAMPLE_QUESTION, expression, TOP_K)
๋ฆฌ์์ค ๋ฐ ์ถ๊ฐ ์๋ฃ
Milvus ๋น ๋ฅธ ์์ ๊ฐ์ด๋
๋ฐฐ์ด ํ๋ ์ฌ์ฉ | Milvus ๋ฌธ์
https://github.com/milvus-io/pymilvus/blob/2.4/examples/fuzzy_match.py
https://milvus.io/docs/boolean.md#Usage
https://milvus.io/docs/single-vector-search.md#Filtered-search
๊ณ์ ์ฝ๊ธฐ

Build Multimodal Search for 3D Assets with Tripo and Zilliz Cloud
Generate 3D assets with Tripo, then search them by text, image, and metadata with multimodal embeddings and Zilliz Cloud.

Data Deduplication at Trillion Scale: How to Solve the Biggest Bottleneck of LLM Training
Explore how MinHash LSH and Milvus handle data deduplication at the trillion-scale level, solving key bottlenecks in LLM training for improved AI model performance.

Zilliz Cloud Delivers Better Performance and Lower Costs with Arm Neoverse-based AWS Graviton
Zilliz Cloud adopts Arm-based AWS Graviton3 CPUs to cut costs, speed up AI vector search, and power billion-scale RAG and semantic search workloads.



