Dear,
We are studying in history field and trying to spot/extract ancient references cited/mentioned in papers in PDF. Please see the samples I have attached. In the sample PDFs, the references are annotated (highlighted yellow).
As you will see, they don't have that much structure and quite different than a regular reference showing up in a scientific paper.
They show up in parts other than full text such as footnotes and abstracts.
So, I could not be sure if creating my train set with createTraining and delete wrong tags and create mines based on the annotations make sense? If I follow this approach, would it be sort of transfer learning as a result the full text model will be improved regarding catching such references we are interested? If so, I guess it would not work since the new data have totally different structure or even no structure? Plus, all the tags aside from ref tags are irrelevant to us, so they have to be deleted. But on the other hand, CRF could still spot them correctly thanks to newly created feature functions?
If createTraining approach does not make sense, I can use "createTrainingBlank" to create train set. As a result, I can only tag ref based on the annotations even including abstracts and footnotes. However, this time I could not be sure what model I should train? Would it be full text model with only ref and p tags (by changing file name from blank.tei.xml to fulltext.tei.xml)? Additionally, if this is the correct approach, I guess we will need way more training sets to be able get satisfactory results?
Many thanks in advance for your guidance!
Pdf_FilenamePeeters_3289944_Annotated.pdf
Pdf_FilenamePeeters_3289945_Annotated.pdf
Pdf_FilenamePeeters_3289948_Annotated.pdf
Pdf_FilenamePeeters_3289953_Annotated.pdf
Pdf_FilenamePeeters_3289954_Annotated.pdf
Dear,
We are studying in history field and trying to spot/extract ancient references cited/mentioned in papers in PDF. Please see the samples I have attached. In the sample PDFs, the references are annotated (highlighted yellow).
As you will see, they don't have that much structure and quite different than a regular reference showing up in a scientific paper.
They show up in parts other than full text such as footnotes and abstracts.
So, I could not be sure if creating my train set with createTraining and delete wrong tags and create mines based on the annotations make sense? If I follow this approach, would it be sort of transfer learning as a result the full text model will be improved regarding catching such references we are interested? If so, I guess it would not work since the new data have totally different structure or even no structure? Plus, all the tags aside from ref tags are irrelevant to us, so they have to be deleted. But on the other hand, CRF could still spot them correctly thanks to newly created feature functions?
If createTraining approach does not make sense, I can use "createTrainingBlank" to create train set. As a result, I can only tag ref based on the annotations even including abstracts and footnotes. However, this time I could not be sure what model I should train? Would it be full text model with only ref and p tags (by changing file name from blank.tei.xml to fulltext.tei.xml)? Additionally, if this is the correct approach, I guess we will need way more training sets to be able get satisfactory results?
Many thanks in advance for your guidance!
Pdf_FilenamePeeters_3289944_Annotated.pdf
Pdf_FilenamePeeters_3289945_Annotated.pdf
Pdf_FilenamePeeters_3289948_Annotated.pdf
Pdf_FilenamePeeters_3289953_Annotated.pdf
Pdf_FilenamePeeters_3289954_Annotated.pdf