LlamaIndex RAG with SovereignEG

Goal: ask questions about your own PDFs and text files (RAG = Retrieval-Augmented Generation).

Time: 15 minutes.

LlamaIndex needs two SovereignEG models:

  • a chat model to write answers → SovereignEG/Qwen3.8-27B-FP8
  • an embedding model to turn text into vectors for search → embeddinggemma-300m

Both are on the same base URL (https://backend.sovereigneg.com/v1) and key.

Step 1 — Install

pip install llama-index llama-index-llms-openai-like llama-index-embeddings-openai-like llama-index-readers-file

Why these four?

PackageJob
llama-indexCore: document loading, index, query engine
llama-index-llms-openai-likeThe OpenAILike chat class. The plain OpenAI class only accepts official OpenAI model names and rejects SovereignEG ids
llama-index-embeddings-openai-likeThe OpenAILikeEmbedding class — same reason, for the embedding model
llama-index-readers-fileLets SimpleDirectoryReader read PDF, DOCX, CSV, etc. (installs pypdf)

Do not name your script llama_index.py. Python will import your file instead of the library and fail with No module named 'llama_index.core'.

Step 2 — Put some files in a data folder

my_rag/
├── rag_query.py
└── data/
    ├── report.pdf
    └── notes.txt

No files handy? Make a test PDF:

pip install fpdf2
from fpdf import FPDF
pdf = FPDF(); pdf.add_page(); pdf.set_font("Helvetica", size=12)
pdf.multi_cell(0, 8, "BlueRiver Support launched an AI chatbot in March 2025. "
                     "It answers 55 percent of simple questions. Revenue grew 18 percent.")
pdf.output("data/report.pdf")

Step 3 — The RAG script

Save as rag_query.py:

import os
import sys
 
from llama_index.core import SimpleDirectoryReader, VectorStoreIndex, Settings
from llama_index.llms.openai_like import OpenAILike
from llama_index.embeddings.openai_like import OpenAILikeEmbedding
 
api_key = os.environ["SOVEREIGNEG_API_KEY"]
base_url = "https://backend.sovereigneg.com/v1"
 
# 1. Chat model — OpenAILike accepts any model name
Settings.llm = OpenAILike(
    api_base=base_url,
    api_key=api_key,
    model="SovereignEG/Qwen3.8-27B-FP8",
    is_chat_model=True,
    context_window=131072,
)
 
# 2. Embedding model — must be an embedding model, not a chat model
Settings.embed_model = OpenAILikeEmbedding(
    api_base=base_url,
    api_key=api_key,
    model_name="embeddinggemma-300m",
    embed_batch_size=10,
)
 
# 3. Load files, build the index, ask a question
data_dir = os.path.join(os.path.dirname(os.path.abspath(__file__)), "data")
if not os.path.isdir(data_dir) or not os.listdir(data_dir):
    print(f"Put some files in {data_dir} first.")
    sys.exit(1)
 
documents = SimpleDirectoryReader(data_dir).load_data()
index = VectorStoreIndex.from_documents(documents)
 
query_engine = index.as_query_engine()
response = query_engine.query("Summarize the main points of my documents.")
 
print("\n--- Response ---")
print(response)
python rag_query.py

What happens:

  1. SimpleDirectoryReader reads every file in data/ and splits it into chunks.
  2. Each chunk is sent to /v1/embeddings once and stored in memory.
  3. Your question is embedded, the closest chunks are found, and they are sent to the chat model with your question. The model answers from your files.

Step 4 — Chat mode (follow-up questions)

chat_engine = index.as_chat_engine(chat_mode="condense_plus_context")
print(chat_engine.chat("What was launched in March 2025?"))
print(chat_engine.chat("How much did it save?"))   # remembers the first question

Step 5 — Save the index (skip re-embedding next time)

from llama_index.core import StorageContext, load_index_from_storage
 
index.storage_context.persist("storage")                  # after building
storage = StorageContext.from_defaults(persist_dir="storage")
index = load_index_from_storage(storage)                  # on later runs

Step 6 — See which chunks were used

response = query_engine.query("What is the budget?")
for node in response.source_nodes:
    print(round(node.score, 3), node.metadata.get("file_name"), node.text[:80])

Choosing an embedding model

Any embedding model from /v1/models works. Copy its id into model_name.

Troubleshooting

ProblemFix
Unknown model '...'You used llama_index.llms.openai.OpenAI. Switch to OpenAILike
'...' is not a valid OpenAIEmbeddingModelTypeOpenAIEmbedding only knows OpenAI names. Use OpenAILikeEmbedding
No module named 'llama_index.core'Your script is named llama_index.py — rename it and delete __pycache__
PDF loads with no textScanned PDF (images). Run OCR first, or use a text export
Slow indexingNormal for big folders; persist the index (Step 5) so it is built once

Next: Open WebUI