Internet LIS 4488 Network Administration for the Information Professional 2U 4U 5U 4U 4U 5U 1 U 41U U 1U 4U Cloud Provider Week 12: Troubleshooting PAGE 1 Agenda General Course Housekeeping Week 11 Vendor Labs Textbook Chapter 9: Troubleshooting Tie-in with Design and Buildout group projects Testable material covered PAGE 2 Textbook and Learning Materials ▪ CompTIA Cloud+ Study Guide (2021) by Ben Piper ISBN: 9781119810865 ▪ LinkedIn Learning Course Materials ▪ Vendor-based Cloud Training PAGE 3 Week 12 Testable Topics Chapter 9: Troubleshooting I ▪ Incident types and prioritization ▪ Incident management and logging ▪ Templates ▪ Time synchronization ▪ Troubleshooting workflows ▪ Capacity Issues and Boundaries ▪ Automation and Orchestration PAGE 4 Cloud Troubleshooting Cloud Troubleshooting A key function of network administrators is troubleshooting. Troubleshooting is mostly a science, but also an art: both can be honed with experience, but there are foundational principles on which one can reliably and effectively begin. A first step is to have a deep understanding of the components and interrelationships between the components. A second is to have a set of tools and methods. PAGE 5 Cloud Troubleshooting To develop troubleshooting expertise, an administrator must expect to engage in continual learning concerning parts of the network and the concepts associated with them – including parts for which the administrator might not be directly responsible. Cloud Troubleshooting Documentation and communication remains critically important. A system may suddenly fail to operate simply because one party decided not to continue paying for an undocumented, uncommunicated element on which a system depends, like a DNS domain name or a certificate or an MFA service. PAGE 6 Cloud Troubleshooting To develop troubleshooting expertise, an administrator should engage in routine what if scenarios. Cloud Troubleshooting “99.9% of our users use VPN-1. If VPN-7 goes down, does that mean that we’re 99.9% functional? Are we good?” “Somewhat, except that some of our executives use VPN-7 because they prefer an endpoint that doesn’t work with VPN-1. Consider the consequences of that: what if VPN-7 goes down while they’re in Atlanta for the annual shareholders’ meeting?” PAGE 7 Cloud Troubleshooting: Incidents Cloud Troubleshooting An incident is any event that poses a threat to normal operations. An incident can involve user error, or a system fault or service failure within or beyond the administrator’s control. Incident Management An incident can be security-related, mandating additional responses, actions, and follow-ups. PAGE 8 Cloud Troubleshooting: Incidents Automated systems can be invaluable in managing and mitigating incidents. Be mindful that automated systems may also be the cause of incidents, once responses are triggered in milliseconds and sometimes by the thousands. Cloud Troubleshooting Incident Management It is extremely important to understand both the nature of incidents, the people amongst whom these incidents occur, the systems involved and their interdependencies, the rules by which incident responses are launched, and their specific effects. PAGE 9 Cloud Troubleshooting: Incidents Cloud Troubleshooting Under certain circumstances, you may have no access to certain parts of infrastructure, e.g. ISP routers upstream of your enterprise edge, or the layers managed by your cloud provider. Incident Management It is wise to have organized in advance mechanisms by which you can contact those providers, and have a very clear understanding as to who has responsibility for what. PAGE 10 Cloud Troubleshooting Baselines traceroutes PAGE 11 Cloud Troubleshooting: Interoperability Interoperability refers to the ability of elements of disparate systems to work with one another, be they of different operating systems or cloud platforms. Cloud Troubleshooting Interoperability This may be a relative term, as processes can be put in place to convert streams, formats, and messages, or they may simply be compatible via the use of mutually compatible forms and structures. Some vendors tout the ability to work with specific other vendors’ systems and products. Open standards help ensure interoperability. PAGE 12 Cloud Troubleshooting: Interoperability When troubleshooting flows across disparate systems, consider whether interoperability might be an issue. Consider how and when any changes were made, how long the system has been in operation, and review histories of issues and concerns. Cloud Troubleshooting Interoperability Review configurations, logs and service tickets. Consider seasonality: a migrated database might have a process that it triggered only on rare occasions, or during certain periods or activities. Regular communication between stakeholders is critical. PAGE 13 Cloud Troubleshooting: Interoperability Cloud Troubleshooting Before, after, and during cloud migrations, review in Detail comparisons between the way that Cloud Vendor A does things, and the way that Cloud Vendor B does them. Interoperability The vendors might provide insight on this – or you might need to seek the advice or technical acumen of consultants. PAGE 14 Cloud Troubleshooting: Connectivity Cloud Troubleshooting Pay close attention to network connectivity between clouds, business premise and other datacenters, and users, and who is responsible for what. Connectivity In networking, open standards help avoid surprises. Monitor dashboards continually. PAGE 15 Cloud Troubleshooting: Licensing Cloud Troubleshooting Licensing is pervasive across all things I/T, from operating systems to applications to devices to services. Licensing Subscription models have become the norm. In days past, organizations worked to ensure compliance with license terms, to avoid audits and lawsuits. They still do, but system elements increasingly “phone home” to the vendor, reporting compliance and may simply stop working if licenses expire, and subscriptions are not renewed. Different vendors handle this in different ways. PAGE 16 Cloud Troubleshooting: Licensing Some vendors allow grace periods to renew subscriptions. Some disable or limit features upon expiration. Some simply annoy the users with reminders to renew. Cloud Troubleshooting Licensing Some vendors will work with customers and extend dates. Some vendors will simply disable services upon expiry. Some system elements will continue to operate, but may become unmanageable or reject updates or telemetry. The bottom line is that it’s extremely important that admins be aware of expiration dates and conditions, contrasts and TOS, who’s responsible for what, procurement status, and system configuration and behavior. It is the admin’s job to know! PAGE 17 Cloud Troubleshooting: Licensing Licensing must be documented and reviewed. Licensing options may vary even within a single vendor and include Cloud Troubleshooting Licensing • total number of users within an organization • total number of concurrent users • total number of connections • other usage metrics License capacity should be planned well in advance of use. PAGE 18 Cloud Troubleshooting: TLS Certificates Cloud Troubleshooting Many web sites, applications, and authentication systems rely on valid and current TLS certificates. If those certificates expire, the functions relying on them break. TLS Certificates It is imperative that certificate expiration dates be tracked to allow maintenance windows if they are generated internally, and to allow that and procurement time if they are purchased. DNS Domain names may also require renewal. PAGE 19 Cloud Troubleshooting: Networking Cloud Troubleshooting While moving applications to the cloud may reduce the need to maintain physical networking hardware and their operating systems, connectivity still must be documented, monitored, baselined, optimized, and occasionally troubleshooted. Networking Cloud ingress and egress points, including Internet and between other cloud and on-premise networks must work efficiently and reliably. PAGE 20 Cloud Management: Baselines Cloud Management It is important to monitor and baseline metrics pertaining to the network. Baselines Bandwidth is the theoretical maximum data transfer capacity of a network. Network Throughput is how much data actually gets through a channel in a given period. Congestion, latencies, and errors/packet loss all work to diminish throughput. Capacity is an end-to-end metric for available bandwidth. Be mindful of bottlenecks in network pathways. PAGE 21 Cloud Management: Baselines Bandwidth is the theoretical maximum data transfer capacity of a network. Packet loss: this is often a function of errors, packets being discarded by devices, or other factors. Cloud Management Baselines Network Latency: servers and other network devices take time to process packets. Some applications are more sensitive than others to packet loss and high latency: e-mail, for example, tends to be highly tolerant, while real-time voice and video tend not to be. Jitter is caused by variable latencies, causing buffering, packets to arrive out of order, and wreaking havoc on voice and video. PAGE 22 Cloud Management Baselines Validation PAGE 23 On-premise Users On-premises Datacenter LAN, MAN, or WAN Off-premise Users Internet External application on Internet Maintain detailed and comprehensive physical and logical diagrams and other documentation pertaining to your enterprise and be sure to include all parts of your systems, internal and external. Include contact information, including site and circuit IDs. Internet Customers Cloud Provider Connectivity CompTIA’s Network+ curriculum covers network troubleshooting in more detail. PAGE 24 Cloud Management: Metrics Remember to correlate network metrics with others: • availability: percentage of uptime • bandwidth utilization: percentage of available bandwidth used • database utilization: database activity in actions per time period • disk capacity: percentage of total disk capacity used • disk utilization: I/O operations per second or other metrics • processor utilization: percentage of available processor used • web server utilization: web activity in actions per time period Cloud Management Baselines PAGE 25 Cloud Management: Baselines and Troubleshooting Baselines need to be validated to ensure that they reflect reality. For example, do they conform with expected usage patterns for particular parts of the year, e.g. student registration or online shopping. Cloud Management Baselines Use in troubleshooting Load simulation and testing can provide useful information. Bottlenecks and other potential root causes can be detected and corrected before they become problems in production. Baselines should be collected before and after upgrades and changes. PAGE 26 Cloud Management: Resource Contention Cloud Management Trending data collected for metrics can be extremely useful in troubleshooting resource contention and starvation issues. These metrics include Baselines Resource Contention • CPU utilization • I/O and other storage-related activity • network bandwidth utilization • web hits and other web activity • application-related activity • database hits and other database activity PAGE 27 Cloud Management: Baselines Cloud Management Internet Baselines Trending public tier: web server Resource Contention second tier: application server Tactics for mitigating resource contention issues include third tier: database server throttling the application horizontal scaling autoscaling Resource contention might occur at any tier of an application stack. Vertical scaling may require downtime, though some instances might support dynamic allocation. PAGE 28 Cloud Troubleshooting: “War rooms” Cloud Troubleshooting A “war room” is a term used to describe a collaborative session and environment where different groups – technical, managerial, application owner – may meet in real time to assist with troubleshooting and executing changes. “War Rooms” They may include vendors and providers, internal technical staff, managers to authorize ad hoc actions, subject matter experts (e.g. business unit leads using applications) and other stakeholders. PAGE 29 Cloud Troubleshooting: Incident Prioritization Incidents are events that affect the enterprise. Cloud Troubleshooting Incidents have different priority categorization and priority levels: impact and urgency. Incident Prioritization Impact involves to the damage and/or disruption caused, e.g. who and what are down. An air traffic control system incident obviously has higher impact than a malfunction causing the airport vending machines to shut down. Urgency refers to the speed at which an incident must be resolved. Understanding one's own organization and its missions, objectives, and workflows is crucial – and must be documented. PAGE 30 Cloud Troubleshooting: Incident Response Documentation Cloud Troubleshooting Procedures for incident response and handling must be documented! Incident Response Documentation • Auditors demand documentation • Turnover is a reality • Experts need and should take vacations • • The expert might not know what state of mind exists during an incident, e.g. distraction, duress, sleepdeprivation Configurations and dependencies also change, so incident response procedures must be periodically reviewed and revised. PAGE 31 Cloud Troubleshooting: Call Trees Cloud Troubleshooting When an incident occurs, who do you notify to assist in the response? Who do you notify to inform of the incident to mitigate impact? Call Trees A call tree is a document that lists entities and persons who must be notified in the event of different types and classes of incidents. This should include multiple contact channels, e.g. phone number, e-mail, and other notification means. Responsibilities for answering should be worked out in advance, e.g. on-call schedules and confirmation of receipt. PAGE 32 Cloud Troubleshooting: Tabletop Exercises Incident response must be planned and rehearsed. Firefighters have an expression: Don’t train until you get it right – train until you can’t get it wrong.” Cloud Troubleshooting Tabletop Exercises Drills are imperative. One such drill is a tabletop exercise. Participants game out scenarios and responses and play out roles as if the incident was real, discussing the strengths and weaknesses of various approaches. Tabletop exercises are a cost-effective means of developing and reviewing incident response plans, evolving documentation, and raising understanding. PAGE 33 Cloud Troubleshooting: Automated Incident Handling Cloud Troubleshooting Some monitoring and logging systems will automatically generate tickets in the organizations incident ticketing system, typically including name and contact information for the party affected, and the name and category of the incident. Automated Incident Handling It is important to understand and document in advance which types and origins of events would affect what, and to what severity. PAGE 34 Cloud Troubleshooting: Time Synchronization Incidents are precipitated by and associated with events. Events are things that happened. Things that happen occur at exact times. Cloud Troubleshooting Time Synchronization To properly correlate events, times need to be precise. For time to be precise, devices and systems need to be precisely and accurately time-synchronized. Network Time Protocol (NTP) is a network protocol that references one or more atomic clocks connected to the Internet: the true time is the same everywhere. In most enterprises, a few internal devices or servers reference Internet NTP sources directly: most in turn reference those. PAGE 35 Cloud Troubleshooting: Time Synchronization Incidents are precipitated by and associated with events. Events are things that happened. Things that happen occur at exact times. Cloud Troubleshooting Time Synchronization To properly correlate events, times need to be precise. For time to be precise, devices and systems need to be precisely and accurately time-synchronized. Network Time Protocol (NTP) is a network protocol that references one or more atomic clocks connected to the Internet: the true time is the same everywhere. In most enterprises, a few internal devices or servers reference Internet NTP sources directly: most in turn reference those. PAGE 36 Cloud Troubleshooting: Cloud Capacity Cloud deployments tend to be extremely reliable and resilient – issues do often arise pertaining to how much capacity has been planned for and allocated to Specific deployments. Cloud Troubleshooting Cloud Capacity If resources are depleted, application and system performance will suffer, and at some point fail. Cloud providers are held to SLAs, and do publish maximum cloud capacities, but it is typically the cloud customer who is held to allocating sufficient resources : if you didn’t purchase them, you might not have them until you do. PAGE 37 Cloud Troubleshooting: Cloud Capacity Cloud customers are billed for resource consumption. Cloud Troubleshooting This includes Cloud Capacity • compute resources e.g. CPU utilization • storage utilization • API calls per period of time, e.g. database or storage writes • bandwidth utilization: inward and/or outward depending on the service and provider • batch job scheduling Applications can be throttled to limit CPU and bandwidth utilization, but can crash if CPUs become overtaxed. Applications and servers will likely crash if storage is depleted. PAGE 38 Cloud Troubleshooting: Cloud Capacity Plan also your internal IP address spaces, even if you are using private IP address ranges. These ranges are usually expressed in CIDR blocks, and are very difficult to change if you run out of IP addresses from not allocating properly. Cloud Troubleshooting Cloud Capacity Private IP address blocks used for cloud subnets are 10.0.0.0 /8 24-bit block 172.16.0.0 /”16” 20-bit block 192.168.0.0 /16 16-bit block up to 16,777,216 IP addresses up to 1,048,576 IP addresses up to 65,536 IP addresses Be mindful of the fact that cloud providers often require separate /24 blocks for each subnet, and that they typically reserve a number of addresses on each subnet for gateways, and DNS and other servers. Divide by 256 to get a rough idea of how many /24 subnets you can get out of each private IP address block. PAGE 39 Cloud Troubleshooting: Orchestration Much of the cloud is abstracted beyond the customer’s direct purview and observation, with it being more or less depending on the –aaS model selected and the responsibility you assume. Cloud Troubleshooting Orchestration Some providers provide limited dashboard capability to inform customers of system states behind those levels of abstraction. It may be necessary to engage the cloud provider technical support staff should suspected problems arise. Online user forums can also be sources of useful information. PAGE 40 Cloud Troubleshooting: Orchestration Cloud Troubleshooting Issues may vary. Common problems include Orchestration • Issues with process and workflow • Issues with accounts and their permissions and roles • Issues with change and configuration management • Issues with patches, upgrades, and versions • Issues with network access control • Issues with naming, IP address change, and DNS • Issues with datacenter and region • Insufficient testing and fit to deployment model PAGE 41 Economics Word-of-the-Week Economics WoW A key difference between cloud and on-premise involves Total Cost of Ownership (TCO) . Variable costs change with units of production. Fixed costs: not so much. PAGE 42 Cloud Vendors Amazon Web Services (AWS) DigitalOcean Cloud Vendors Google Cloud Platform (GCP) IBM Cloud Microsoft Azure Oracle Cloud Platform (OCP) Note: these are just examples of cloud providers, and this is not a complete list. PAGE 43 Here is what we learned ▪ Incident types and prioritization ▪ Incident management and logging ▪ Templates Week 12 Summary ▪ Time synchronization ▪ Troubleshooting workflows ▪ Capacity Issues and Boundaries ▪ Automation and Orchestration First Skill Second Skill Third Skill Conclusion PAGE 44
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )