I think the extracted content should be in the Content field. See https://docs.optimizely.com/graph/docs/text-extraction-from-media
Not the _fulltext but it should be included in CMS12, I believe for pdfs, docx, xls, xlsz, txt
Thanks! I must be missing something because I get no results using:
{
GenericMedia(limit: 100, where: { Content: { exist: true } }) {
items {
_fulltext
Content
ContentType
}
}
}
I've now seen text extraction work in my local 13 PaaS solution... It goes into _fulltext same way as I confirmed later with SaaS.
There's ton of files in this setup that I was hoping text could be extracted from, but where it wasn't, that fooled me. A simpler PDF (exported from Word) works though.
I will try in 12 with this file shortly as well...
I tested this with both CMS 12 and CMS 13.
In CMS 12, extraction worked for the PDF/DOCX files I tested. The extracted text was available in Content, _fulltext and _json, and Content: { match: "..." } successfully found text from the documents.
One thing I noticed is that Content: { exist: true/false } does not seem to behave reliably for checking whether extracted text exists. I could get exist: false results even when Content was populated. I also found that a normal string field with an empty value ("") still matches exist: true.
In CMS 13, the schema is different: the Content field is no longer present, but the extracted text is available in _fulltext, and searching it works correctly.
So extraction itself seems to work in both versions, but the schema/representation differs, and on CMS 12 I wouldn't rely on Content.exist to determine whether extraction succeeded.
If so, what is required?
I'm testing with a content type that has pdf extension set. Seeing that _fulltext does get some meta attributes etc but nothing from the text content of the file.
{GenericMedia(
limit:100
where: {
Name: { contains: "test.pdf" }
}
) {
items
{
_fulltext
}
}
}
These are the versions installed: