Introduction
At PiXirus, Artificial Intelligence (AI) and Machine Learning (ML) are at the heart of our digital innovation strategy. From large-scale digitization to eLearning and content management, we design intelligent solutions that preserve knowledge, enhance discovery, and transform access for libraries, archives, museums, and enterprises worldwide.
We are expanding our AI team and are looking for passionate AI Developers who want to shape the future of cultural heritage, knowledge management, and digital learning.
Objective
Our objective is to build a global AI team that can design, implement, and scale AI-powered workflows for:
- Digitization & preservation of vast image / data collections
- OCR & handwriting recognition
- Automated cataloguing, indexing, and metadata creation
- AI-driven document quality control
We welcome applications from anywhere in the world. You can choose to:
- Work remotely from your location, or
- Join us in our offices in Pondicherry (India) or France.
Qualifications
- Bachelor’s / Master’s / PhD in Computer Science, Artificial Intelligence, Data Science, or related fields
- Strong foundation in machine learning, deep learning, and NLP
- Knowledge of document AI, OCR, computer vision, or search optimization is a plus
- Excellent programming skills (Python, familiarity with Java / C++ is a plus)
- Good communication skills and ability to collaborate in global teams
- Language skills: English – essential for global collaboration, French – highly desirable for working with cultural heritage institutions and European clients
Key Skills & Tools
- AI / ML Frameworks: PyTorch, TensorFlow, scikit-learn
- OCR & HWR: Tesseract, Kraken OCR, Calamari, TrOCR
- NLP & Metadata Extraction: Hugging Face Transformers, spaCy, GROBID
- Search & Indexing: ElasticSearch, Solr, FAISS, Pinecone
- Data Pipelines: Airflow, Spark, Pandas, MongoDB / PostgreSQL
- Document Processing: OpenCV, PyPDF2, XML/EPUB toolkits
- MLOps: MLflow, Docker, Kubernetes, CI/CD pipelines
- Generative AI & LLM Tools:
‣ ChatGPT API (OpenAI) – dynamic integration for summarization, translation, Q&A
‣ Google AI Studio (Gemini) – multimodal AI for text, vision, and search
‣ Other GPTs (Claude, LLaMA, Mistral, Falcon) – fine-tuning & domain adaptation
‣ LangChain / LlamaIndex – orchestration of LLMs for pipelines
‣ Vector Databases (Pinecone, Weaviate, Milvus) – for semantic retrieval with LLMs
Roles & Responsibilities
- Develop, train, and optimize AI/ML models for OCR, NLP, and document AI
- Automate cataloguing, indexing, and metadata tagging of cultural heritage collections
- Design AI workflows for digitization, recognition, and quality control
- Enhance search and discovery systems using semantic search & embeddings
- Collaborate with archivists, librarians, eLearning teams, and engineers
- Deploy models into scalable production environments with MLOps best practices
- Document processes and participate in knowledge transfer & training programs
Work Locations
How to Apply
Job Types: Full-time, Permanent
Pay: From ₹15,000.00 per month
Benefits:
Work Location: In person