Skip to content

Platform Infrastructure & Performance

Epic Overview

EPIC: Platform Infrastructure & Performance

DESCRIPTION: Design and implement a robust, scalable, and high-performance infrastructure for the REC Verifiable Credentialing Platform. The infrastructure will be cloud-hosted, support multi-tenancy, ensure high availability, implement disaster recovery mechanisms, and meet performance requirements for response times and concurrent users. The system will be designed to scale both horizontally and vertically as demand increases.

BUSINESS VALUE: Provides a reliable, responsive, and resilient foundation for the platform, ensuring consistent service delivery even during peak usage. The scalable architecture allows the platform to grow with increasing adoption, while the high availability and disaster recovery capabilities minimize downtime and data loss risks, protecting both business continuity and reputation.

STAKEHOLDERS: - REC (Platform Owner) - Staffing Companies - Candidates - Velocity Career Labs (VCL) - Operations and Support Teams

SIZE: Large - Infrastructure is a critical component that requires significant planning and implementation effort.

USER STORIES: - As a platform administrator, I want the solution to be cloud-hosted so that we can leverage cloud scalability and reliability. - As a platform administrator, I want the system to be scalable both horizontally and vertically so that it can handle increasing demand. - As a platform administrator, I want to support a multi-tenanted environment so that we can serve hundreds of staffing companies. - As a platform administrator, I want to ensure 99.9% uptime with self-healing capabilities so that the service remains available to users. - As a platform administrator, I want to implement disaster recovery plans so that the system can be fully restored within 24 hours if needed. - As a platform administrator, I want to implement regular data backups so that data can be retrieved for the last 30 days. - As a staffing company user, I want the user interfaces to be mobile-responsive so that I can access the platform from any device. - As a staffing company user, I want the system to handle high transaction volumes so that operations aren't delayed during peak times. - As a staffing company user, I want response times under 2 seconds for all major actions so that I can work efficiently. - As a candidate, I want the platform to be responsive and reliable so that I can easily manage my credentials.

Implementation Details

The Platform Infrastructure & Performance system includes:

  1. Cloud Infrastructure
  2. AWS-based hosting (preferred for co-location with Velocity Credential Agent)
  3. Infrastructure as Code (IaC) implementation
  4. Container-based deployment
  5. Microservices architecture where appropriate

  6. Scalability

  7. Horizontal scaling through load balancing
  8. Vertical scaling capabilities
  9. Auto-scaling based on demand
  10. Database scaling strategies

  11. Multi-tenancy

  12. Secure data segregation between companies
  13. Tenant-specific configurations
  14. Shared infrastructure with isolated data
  15. Support for hundreds of staffing companies

  16. High Availability

  17. 99.9% uptime SLA
  18. Active-active setup
  19. Self-healing and redundancy measures
  20. Automatic failover within 5 minutes
  21. Geographically distributed resources

  22. Disaster Recovery

  23. Full system restoration within 24 hours
  24. Regular testing of recovery procedures
  25. Documented recovery processes
  26. Recovery point objective (RPO) and recovery time objective (RTO) definitions

  27. Data Backup

  28. Daily backups (minimum)
  29. 30-day retention period
  30. Geographically distributed backup storage
  31. Backup verification and testing

  32. Performance Optimization

  33. Support for 1000-2000 transactions per hour
  34. Support for 1000 concurrent users
  35. Support for 100 concurrent registration requests per minute
  36. Response time under 2 seconds for all major actions
  37. Performance monitoring and optimization

  38. Responsive Design

  39. Mobile-responsive user interfaces
  40. Adaptive layouts for different screen sizes
  41. Touch-friendly interface elements
  42. Consistent experience across devices

Dependencies

  • Cloud provider services (AWS preferred)
  • Infrastructure automation tools
  • Monitoring and alerting systems
  • Database management systems
  • Content delivery networks

Performance Requirements

  • Transaction volume: 1000-2000 transactions per hour
  • Concurrent users: 1000 during peak load times
  • Concurrent registration requests: 100 per minute
  • Response time: Under 2 seconds for all major user actions
  • Uptime: 99.9% SLA
  • Recovery time: Full system restoration within 24 hours