I receive a lot of 'File Too Large' warnings and 'Extraction failed ' messages when indexing is running, and I see that many of my files are not being indexed.
What is the expected maximum size of file that can be indexed?
Will failures be retried every time File-Brain is instantiated?
(which would apparently have indexing constantly running, since it fails at least as much as it succeeds. It has currently succeeded on 6 documents out of 14 attempted, and only after several restarts. This is after running for two hours against a corpus of 50K documents waiting to be indexed. Clearly, this will be non-viable unless there is something I can do to improve its performance and success rate.)
I receive a lot of 'File Too Large' warnings and 'Extraction failed ' messages when indexing is running, and I see that many of my files are not being indexed.
What is the expected maximum size of file that can be indexed?
Will failures be retried every time File-Brain is instantiated?
(which would apparently have indexing constantly running, since it fails at least as much as it succeeds. It has currently succeeded on 6 documents out of 14 attempted, and only after several restarts. This is after running for two hours against a corpus of 50K documents waiting to be indexed. Clearly, this will be non-viable unless there is something I can do to improve its performance and success rate.)