Can text-embedding-3-small answer questions directly?
No. It converts questions or materials into vectors for applications to compare semantic similarity. Knowledge-base Q&A typically uses it first to retrieve relevant passages, then passes them to a response model to compose an answer; calling only the embedding API returns vectors, not explanations, summaries, or conversational replies.
How does it differ from text-embedding-ada-002?
It is a third-generation embedding model, not an alias for ada-002. Official release benchmarks show that it achieves higher average scores on multilingual retrieval and English tasks, and it supports shortening vectors through dimensions. When migrating, rebuild document vectors to avoid comparing results from the two models in the same space.
When should text-embedding-3-large be chosen?
When retrieval precision matters more than lightweight operation, it is worth comparing large. It scored higher in the official benchmarks released at the same time, but the choice should still be based on your own query set and document repository. You can first establish a small baseline, then evaluate whether large improves retrieval of relevant passages for key questions.
How do I submit batches and match returned results?
Submit model and input to POST /v1/embeddings. input can use batches of text arrays or token arrays, with a maximum of 2048 items. Each item in the returned data contains index and embedding; use index to align with the original input, and check usage for token consumption.
What do dimensions and encoding_format control respectively?
dimensions controls vector shortening; if unspecified, the full dimensions are output. encoding_format controls the returned representation, with float or base64 available and float as the default. The former requires evaluating retrieval performance, while the latter is for adapting to data-processing methods; encoding choices should not be understood as model accuracy levels.