Map the system and the symptoms
Identify the services, data locations, restarts and errors users encounter. Collect versions, dependencies, logs and observations within the agreed access. Separate symptoms from causes: high load may result from a queue rather than an undersized server. For intermittent problems, consider safe additional monitoring. The outcome should be a clear system picture and supported findings, not a list of guesses demanding immediate replacement of every component without evidence.
Keep security checks within an agreed scope
User permissions, secret storage, administrative exposure and incoming-data handling deserve specific attention. Agree on scope and acceptable load in advance. Active checks that could affect availability should not appear unexpectedly in production. Findings need reproducible evidence and an explanation of impact without exposing passwords in the report. A review is not a certificate of absolute security; it records discovered issues and the boundaries of what was examined at that time.
Turn findings into a workable sequence
Not every issue has the same urgency. Distinguish data-loss risk, unauthorised access, broken customer journeys and improvements that can wait. Important changes need verification and rollback planning. Agree on implementation separately from the review so scope and impact remain clear. Include older bots, websites and scheduled tasks in the dependency picture; improving one application should not leave the owner with an unexpected failure in a neighbouring service.
Two useful questions
Does a review require taking the service offline?
No. Much of the picture can be gathered while the service runs. Any check requiring load or intervention needs its conditions agreed in advance.
Can we commission only the most important fixes?
Yes. Priorities support an informed choice. Start with risks to data and core journeys, then plan other improvements separately.