LlamaIndex RAG with SovereignEG
Goal: ask questions about your own PDFs and text files (RAG = Retrieval-Augmented Generation).
Time: 15 minutes.
LlamaIndex needs two SovereignEG models:
- a chat model to write answers →
SovereignEG/Qwen3.8-27B-FP8 - an embedding model to turn text into vectors for search →
embeddinggemma-300m
Both are on the same base URL (https://backend.sovereigneg.com/v1) and key.
Step 1 — Install
pip install llama-index llama-index-llms-openai-like llama-index-embeddings-openai-like llama-index-readers-fileWhy these four?
| Package | Job |
|---|---|
llama-index | Core: document loading, index, query engine |
llama-index-llms-openai-like | The OpenAILike chat class. The plain OpenAI class only accepts official OpenAI model names and rejects SovereignEG ids |
llama-index-embeddings-openai-like | The OpenAILikeEmbedding class — same reason, for the embedding model |
llama-index-readers-file | Lets SimpleDirectoryReader read PDF, DOCX, CSV, etc. (installs pypdf) |
Do not name your script
llama_index.py. Python will import your file instead of the library and fail withNo module named 'llama_index.core'.
Step 2 — Put some files in a data folder
my_rag/
├── rag_query.py
└── data/
├── report.pdf
└── notes.txtNo files handy? Make a test PDF:
pip install fpdf2from fpdf import FPDF
pdf = FPDF(); pdf.add_page(); pdf.set_font("Helvetica", size=12)
pdf.multi_cell(0, 8, "BlueRiver Support launched an AI chatbot in March 2025. "
"It answers 55 percent of simple questions. Revenue grew 18 percent.")
pdf.output("data/report.pdf")Step 3 — The RAG script
Save as rag_query.py:
import os
import sys
from llama_index.core import SimpleDirectoryReader, VectorStoreIndex, Settings
from llama_index.llms.openai_like import OpenAILike
from llama_index.embeddings.openai_like import OpenAILikeEmbedding
api_key = os.environ["SOVEREIGNEG_API_KEY"]
base_url = "https://backend.sovereigneg.com/v1"
# 1. Chat model — OpenAILike accepts any model name
Settings.llm = OpenAILike(
api_base=base_url,
api_key=api_key,
model="SovereignEG/Qwen3.8-27B-FP8",
is_chat_model=True,
context_window=131072,
)
# 2. Embedding model — must be an embedding model, not a chat model
Settings.embed_model = OpenAILikeEmbedding(
api_base=base_url,
api_key=api_key,
model_name="embeddinggemma-300m",
embed_batch_size=10,
)
# 3. Load files, build the index, ask a question
data_dir = os.path.join(os.path.dirname(os.path.abspath(__file__)), "data")
if not os.path.isdir(data_dir) or not os.listdir(data_dir):
print(f"Put some files in {data_dir} first.")
sys.exit(1)
documents = SimpleDirectoryReader(data_dir).load_data()
index = VectorStoreIndex.from_documents(documents)
query_engine = index.as_query_engine()
response = query_engine.query("Summarize the main points of my documents.")
print("\n--- Response ---")
print(response)python rag_query.pyWhat happens:
SimpleDirectoryReaderreads every file indata/and splits it into chunks.- Each chunk is sent to
/v1/embeddingsonce and stored in memory. - Your question is embedded, the closest chunks are found, and they are sent to the chat model with your question. The model answers from your files.
Step 4 — Chat mode (follow-up questions)
chat_engine = index.as_chat_engine(chat_mode="condense_plus_context")
print(chat_engine.chat("What was launched in March 2025?"))
print(chat_engine.chat("How much did it save?")) # remembers the first questionStep 5 — Save the index (skip re-embedding next time)
from llama_index.core import StorageContext, load_index_from_storage
index.storage_context.persist("storage") # after building
storage = StorageContext.from_defaults(persist_dir="storage")
index = load_index_from_storage(storage) # on later runsStep 6 — See which chunks were used
response = query_engine.query("What is the budget?")
for node in response.source_nodes:
print(round(node.score, 3), node.metadata.get("file_name"), node.text[:80])Choosing an embedding model
Any embedding model from /v1/models works. Copy its id into model_name.
Troubleshooting
| Problem | Fix |
|---|---|
Unknown model '...' | You used llama_index.llms.openai.OpenAI. Switch to OpenAILike |
'...' is not a valid OpenAIEmbeddingModelType | OpenAIEmbedding only knows OpenAI names. Use OpenAILikeEmbedding |
No module named 'llama_index.core' | Your script is named llama_index.py — rename it and delete __pycache__ |
| PDF loads with no text | Scanned PDF (images). Run OCR first, or use a text export |
| Slow indexing | Normal for big folders; persist the index (Step 5) so it is built once |
Next: Open WebUI