kempersc/glam - Forgejo: Beyond coding. We Forge.

603 commits 2 branches 0 tags 2.5 GiB

Author	SHA1	Message	Date
kempersc	e5a532a8bc	Add comprehensive tests for NLP institution extraction and RDF partnership integration - Introduced `test_nlp_extractor.py` with unit tests for the InstitutionExtractor, covering various extraction patterns (ISIL, Wikidata, VIAF, city names) and ensuring proper classification of institutions (museum, library, archive). - Added tests for extracted entities and result handling to validate the extraction process. - Created `test_partnership_rdf_integration.py` to validate the end-to-end process of extracting partnerships from a conversation and exporting them to RDF format. - Implemented tests for temporal properties in partnerships and ensured compliance with W3C Organization Ontology patterns. - Verified that extracted partnerships are correctly linked with PROV-O provenance metadata.	2025-11-19 23:20:47 +01:00
kempersc	5e9f54bd91	Deduplicate Brazilian institutions (212→121) - Merged 91 duplicate Brazilian institution records - Improved Wikidata coverage from 26.4% to 38.8% (+12.4pp) - Created intelligent merge strategy: - Prefer records with higher confidence scores - Merge locations (prefer most complete) - Combine all unique identifiers - Combine all unique digital platforms - Combine all unique collections - Add provenance notes documenting merges - Create backup before deduplication - Generate comprehensive deduplication report Dataset changes: - Total institutions: 13,502 → 13,411 - Brazilian institutions: 212 → 121 - Coverage: 47/121 institutions with Q-numbers (38.8%)	2025-11-11 22:08:34 +01:00
kempersc	59c99bfb26	Brazil Batch 10: Enrich 8 institutions (26.4% coverage) - Add Wikidata Q-numbers to 8 Brazilian institutions - Coverage: 56/212 institutions (26.4%, +5.6pp gain) - All Q-numbers validated via Wikidata authenticated API - Largest single batch gain yet - Note: Duplicate entries detected, deduplication needed Q-numbers added: - Q10333651 - Museu da Borracha - Q10387829 - UFAC Repository - Q10345196 - Parque Memorial Quilombo dos Palmares - Q1434444 - Teatro Amazonas - Q116921020 - Centro Cultural dos Povos da Amazônia - Q7894381 - UNIFAP - Q16496091 - Arquivo Público do Estado da Bahia - Q56695457 - Museu de Arqueologia e Etnologia da UFPR	2025-11-11 22:05:43 +01:00

Author

SHA1

Message

Date

kempersc

e5a532a8bc

Add comprehensive tests for NLP institution extraction and RDF partnership integration

- Introduced `test_nlp_extractor.py` with unit tests for the InstitutionExtractor, covering various extraction patterns (ISIL, Wikidata, VIAF, city names) and ensuring proper classification of institutions (museum, library, archive).
- Added tests for extracted entities and result handling to validate the extraction process.
- Created `test_partnership_rdf_integration.py` to validate the end-to-end process of extracting partnerships from a conversation and exporting them to RDF format.
- Implemented tests for temporal properties in partnerships and ensured compliance with W3C Organization Ontology patterns.
- Verified that extracted partnerships are correctly linked with PROV-O provenance metadata.

2025-11-19 23:20:47 +01:00

kempersc

5e9f54bd91

Deduplicate Brazilian institutions (212→121)

- Merged 91 duplicate Brazilian institution records
- Improved Wikidata coverage from 26.4% to 38.8% (+12.4pp)
- Created intelligent merge strategy:
  - Prefer records with higher confidence scores
  - Merge locations (prefer most complete)
  - Combine all unique identifiers
  - Combine all unique digital platforms
  - Combine all unique collections
- Add provenance notes documenting merges
- Create backup before deduplication
- Generate comprehensive deduplication report

Dataset changes:
- Total institutions: 13,502 → 13,411
- Brazilian institutions: 212 → 121
- Coverage: 47/121 institutions with Q-numbers (38.8%)

2025-11-11 22:08:34 +01:00

kempersc

59c99bfb26

Brazil Batch 10: Enrich 8 institutions (26.4% coverage)

- Add Wikidata Q-numbers to 8 Brazilian institutions
- Coverage: 56/212 institutions (26.4%, +5.6pp gain)
- All Q-numbers validated via Wikidata authenticated API
- Largest single batch gain yet
- Note: Duplicate entries detected, deduplication needed

Q-numbers added:
- Q10333651 - Museu da Borracha
- Q10387829 - UFAC Repository
- Q10345196 - Parque Memorial Quilombo dos Palmares
- Q1434444 - Teatro Amazonas
- Q116921020 - Centro Cultural dos Povos da Amazônia
- Q7894381 - UNIFAP
- Q16496091 - Arquivo Público do Estado da Bahia
- Q56695457 - Museu de Arqueologia e Etnologia da UFPR

2025-11-11 22:05:43 +01:00

603 commits