CURRENT OPENINGS

Data Scientist

Data Scientist with Bachelor’s degree in Computer Science, Computer Information Systems, Information Technology, or a combination of education and experience equating to the U.S. equivalent of a Bachelor’s degree in one of the aforementioned subjects.

Job Duties and Responsibilities:

  • Collaborate with product owners, data scientists, data engineers, analysts, architects, and business stakeholders to identify high-value data, AI, and analytics opportunities.
  • Translate business requirements into analytical approaches, technical specifications, success metrics, acceptance criteria, and implementation plans.
  • Collect, profile, analyze, and interpret large structured, semi-structured, and unstructured datasets to uncover trends, patterns, anomalies, and business opportunities.
  • Perform exploratory data analysis, statistical analysis, hypothesis testing, correlation analysis, segmentation, and feature engineering.
  • Develop, train, tune, and evaluate machine-learning models for regression, classification, clustering, forecasting, recommendation, optimization, and anomaly detection.
  • Apply advanced statistical and machine-learning techniques, including ensemble methods, time-series modeling, causal inference, and predictive analytics.
  • Select suitable algorithms and evaluate models using precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact metrics.
  • Design experiments and A/B tests to validate hypotheses, measure model effectiveness, and quantify the business impact of data-driven solutions.
  • Develop business dashboards, reports, data visualizations, and self-service analytics solutions that communicate findings to technical and nontechnical stakeholders.
  • Define key performance indicators, build reusable analytical datasets, and perform ad hoc and root-cause analyses to support strategic and operational decisions.
  • Develop end-to-end Retrieval-Augmented Generation pipelines using enterprise data, embeddings, semantic search, hybrid search, reranking, vector databases, and large language models.
  • Apply prompt engineering, grounding, citation validation, structured outputs, evaluation frameworks, and guardrails to improve the accuracy and reliability of Generative AI applications.
  • Design Agentic AI solutions capable of planning, reasoning, tool selection, workflow execution, reflection, and autonomous task completion.
  • Build single-agent, multi-agent, and supervisor-agent workflows using LangGraph, LangChain, LlamaIndex, or comparable Agentic AI frameworks.
  • Integrate AI agents with enterprise APIs, databases, vector stores, documents, cloud services, analytics platforms, and business applications.
  • Implement function calling, tool calling, workflow routing, agent memory, state management, checkpointing, and Model Context Protocol integrations.
  • Develop human-in-the-loop approvals, confidence thresholds, escalation processes, fallback mechanisms, retry policies, and error-recovery controls for agentic workflows.
  • Design scalable data models and analytical schemas using MongoDB, relational databases, cloud data warehouses, data lakes.
  • Develop and maintain ETL/ELT pipelines for data ingestion, cleansing, transformation, validation, normalization, enrichment, and preparation.
  • Build and support batch, streaming, Change Data Capture, and event-driven data pipelines across operational, analytical, and AI systems.
  • Implement data-quality rules, metadata management, lineage, observability, schema-evolution handling, and governance controls.
  • Develop Python-based data and AI services and integrate machine-learning, Generative AI, and Agentic AI capabilities through APIs and reusable components.
  • Package, deploy, version, monitor, and maintain machine-learning models, AI agents, analytical applications, and data pipelines across cloud and enterprise environments.
  • Monitor model accuracy, agent behavior, data drift, model drift, pipeline health, application performance, latency, token consumption, security, and infrastructure costs.
  • Evaluate AI solutions for hallucinations, bias, prompt injection, sensitive-data exposure, unauthorized tool usage, and regulatory risks while documenting findings and coordinating delivery across cross-functional, onshore, and offshore teams.

Technologies Involved / Skills required for the position:

  • Strong programming skills in Python and SQL, including object-oriented programming, data structures, exception handling, debugging, and performance optimization.
  • Experience with data science libraries and machine-learning frameworks such as Pandas, NumPy, Scikit-learn, XGBoost, LightGBM, TensorFlow, and PyTorch.
  • Strong knowledge of exploratory data analysis, statistical analysis, hypothesis testing, feature engineering, causal inference, experiment design, and A/B testing.
  • Experience developing regression, classification, clustering, forecasting, recommendation, optimization, and anomaly-detection models.
  • Knowledge of model-evaluation techniques and metrics, including precision, recall, F1 score, AUC, RMSE, MAE, explainability, fairness, and business-impact measurement.
  • Experience developing dashboards, reports, and self-service analytics using Power BI, Tableau, Looker, Omni, or comparable business-intelligence platforms.
  • Hands-on knowledge of Generative AI, large language models, prompt engineering, embeddings, grounding, structured outputs, tool calling, and guardrails.
  • Experience building RAG solutions using LangChain, LangGraph, LlamaIndex, Hugging Face, or comparable document-processing and AI orchestration frameworks.
  • Knowledge of Agentic AI systems involving reasoning, planning, reflection, workflow routing, tool execution, memory, state management, and autonomous task completion.
  • Experience implementing function calling, tool calling, API integrations, checkpointing, and Model Context Protocol integrations.
  • Knowledge of vector databases and search technologies such as MongoDB Atlas Vector Search, Pinecone and Chroma.
  • Experience integrating AI models through OpenAI, Anthropic, Google Gemini, Vertex AI, AWS Bedrock, Azure OpenAI, or Hugging Face APIs.
  • Knowledge of developing Python-based backend services and REST APIs using FastAPI, Flask, or comparable frameworks.
  • Experience with data-processing and distributed-computing technologies, including Pandas, NumPy, PySpark, SQL, and data-validation frameworks.
  • Experience with relational and NoSQL databases, including PostgreSQL, MySQL, SQL Server, MongoDB, and DynamoDB.
  • Experience with cloud data warehouses, data lakes, or Lakehouse platforms such as BigQuery, Snowflake, Azure Synapse, or Databricks.
  • Experience developing ETL/ELT, batch, streaming, Change Data Capture, and event-driven data pipelines.
  • Knowledge of workflow-orchestration and messaging technologies such as Apache Airflow, Kafka, Google Cloud Pub/Sub, AWS SQS, SNS, and Event Bridge.
  • Knowledge of data-quality rules, metadata management, data lineage, observability, schema evolution, and data-governance controls.
  • Knowledge of deploying AI/ML applications using Docker, Kubernetes, and cloud platforms.
  • Understanding MLOps and LLMOps practices, including model versioning, experiment tracking, CI/CD, automated testing, controlled releases, rollback, retraining, and production monitoring.
  • Knowledge of AI security and responsible AI practices, including hallucination evaluation, bias detection, prompt-injection prevention, sensitive-data protection, explainability, model governance, privacy, and regulatory compliance.
  • Excellent analytical, problem-solving, architecture, code-review, documentation, communication, collaboration, and stakeholder-management skills.

Work location is Portland, ME with required travel to client locations throughout USA.

Rite Pros is an equal opportunity employer (EOE).

Please Mail Resumes to:
Rite Pros, Inc.
565 Congress St, Suite # 305
Portland, ME 04101.

Email: resumes@ritepros.com