# Welcome

Hello, and welcome to the public pages of my cloud devops automation consulting / contracting business.  &#x20;

As you may see from my [linkedin ](https://www.linkedin.com/in/oliver-schoenborn)profile, I have a significant amount of experience in software engineering and devops automation. I would love to discuss how Sentian Cloud Computing Inc could help your business, employer or client (as the case may be!).&#x20;

Since late 2017 I have been focussing on the following:&#x20;

* **Cloud infrastructure provisioning automation**: infrastructure as code using Terraform (also cdktf, Pulumi), Python, bash (increasingly Go)&#x20;
* **Cloud native application deployment and delivery automation**: via CI/CD such as jenkins, spinnaker, github actions, bitbucket pipelines, circleci, gitlab, helm charts, helmfile
* **Kubernetes**: cluster provisioning, management, and securing; helm charts, migration of legacy or docker-compose apps to micro-services in kubernetes, scalability:
* **Serverless** applications: engineering and provisioning "cloud functions" (eg AWS Lambda)

For further details on why you might want to hire me rather than the type of work that can be done via *Sentian Cloud Computing Inc*, my availability, etc, please have a look at [Let's Talk / Recruiters & Co](https://app.gitbook.com/s/-MNs5Wv9Z8JMY1K3-8GZ/~/changes/DbfqtYxiUSZB0aq5zs7W/lets-talk/contracting), where you can also find a link to my calendly.

Sincerely,&#x20;

Oliver Schoenborn\
Cloud DevOps Engineering Contractor / Consultant\
Ontario, Canada


# Why Me / My Business

Some reasons why contracting me through my business might be the right decision for your business, client or employer:

* **Experience**: 4+ years in Cloud DevOps focussed on infrastructure as code and kubernetes; 15+ years before that in software engineering for distributed virtual reality applications.
* **Passion and self-motivation**: I am passionate about cloud devops automation; I stay informed, attend technical meetups regularly, conferences when I can, update my skills with some formal training or workshops once in a while.&#x20;
* **International**: through my business, Sentian Cloud Computing Inc, I can work for your business regardless of its location, with minimal overhead.
* **Low overhead**: no office costs to you, all I need are a few accounts to connect to your cloud infrastructure, CI/CD, communications tools (Slack, Zoom, Teams, Google Meet). I have quality computing hardware where I work.&#x20;
* **Work ethics**: I track all my work time via an app, I stop the counter when I need to shift to non-work related online or offline tasks, I actively attend required meetings, I lookout for opportunities to improve your processes and reliability and cost of your infrastructure and CI/CD
* **Security**: I use virtual machines, business quality password management software, etc to keep contract work completely siloed. &#x20;
* **Agility**: I am happy to do on-site visits for special meetings, to ramp hours down for maintenance periods, and back up to make further improvements, upgrades, etc, and to take part in agile scrum meetings or other Agile processes.
* **Save money**: I engineer automation that just works, and doesn't need constant attention. I do not need all the benefits you offer, keep them for employees who need them!&#x20;
* **Know-how accumulation**: every team I contract for is different, I learn from each one different ways of solving problems, different tools, different concerns, and this gives me a broad view of the devops landscape of tooling and procedures.&#x20;


# Recruiters & Co

To Recruiters, Head Hunters, Talent Seekers, Hiring Managers, etc, wanting to contact me ([Oliver Schoenborn](https://www.linkedin.com/in/oliver-schoenborn/)) about current or potential work opportunities, please bear in mind that without setting boundaries, I would get too many connection and meeting requests, most of which would be a waste of both my time and yours.&#x20;

For this reason, I recommend that you contact me ONLY if you can answer YES to ***all*** of the following regarding this or future opportunities:

* Opportunity focuses on at least 2 of the following :&#x20;
  * cloud infrastructure as code (terraform, pulumi)
  * kubernetes
  * python / Go
* Opportunity involves mostly linux
* Opportunity involves at least one of AWS, Azure, or GCP
* Opportunity is mostly remote (once-in-a-while site visits are ok, like a day or 2 per month)
* Opportunity can be c2c (corp to corp): contract is with my corporation (directly or via a contracting agency eg Toptal, Procom, etc)

Thank you for your understanding!

The following table provides other parameters related to contracting Sentian Cloud Computing Inc to do work for you or your client:&#x20;

***Last updated: June 22, 2022***

<table><thead><tr><th width="258.66134541441824">Aspect</th><th width="467.61821797003057">Constraints</th><th> </th></tr></thead><tbody><tr><td><strong>Availability</strong></td><td><p>As of latest update, my expected availability is </p><ul><li>For 10 hrs / week: November 2022</li><li>For anything > 15 hrs / week: December 2022</li></ul></td><td></td></tr><tr><td><strong>Connection Requests</strong></td><td><p>I accept <a href="https://linkedin.com/in/oliver-schoenborn">LinkedIn connection requests</a> <em><strong>only from people I have met</strong></em>. </p><p></p><p>You are welcome to book a video or audio meeting through my <a href="https://calendly.com/sentian-cloud-computing-inc/15min">Calendly</a>. </p></td><td></td></tr><tr><td><strong>Skills &#x26; Experience</strong></td><td><p>Only contact me about work that involves</p><ul><li>automating cloud infrastructure (AWS, Azure, GCP, IBM, etc) via Terraform, Pulumi, CDK, Docker, Python, Go, bash</li><li>automating application / micro-services deployments via docker, kubernetes, helm charts, python, bash and CI/CD</li><li>setting up, managing, extending (via Golang) and securing kubernetes clusters</li><li>migrating legacy web apps (C++,  Typescript) to kubernetes in Go</li><li>auditing cloud infrastructure and kubernetes for security </li><li>coaching on any of the above</li></ul></td><td></td></tr><tr><td><strong>Permanent Employment?</strong></td><td><p>Contract &#x26; consulting work only, through <em><strong>Sentian Cloud Computing Inc</strong></em>. My business can be hired directly or through other agencies such as Toptal, The AIM Group, GSquad, etc. </p><p></p><p>Rest assured: I have never dropped a client! I finish all work and I have several repeat clients. </p></td><td></td></tr><tr><td><strong>Relocation, On-Site work</strong></td><td>I am happy to visit the workplace when necessary (a couple of times a month), but relocation is not possible. </td><td></td></tr><tr><td><strong>Duration</strong></td><td><p>Weeks, months, a year, open-ended, retainer, hourly, part-time, full time, anything goes. </p><p></p><p>You can end a contract on short notice for whatever reason (funding, change of direction, etc), as long as all work done before I got notice of the termination gets paid.</p></td><td></td></tr><tr><td><strong>Citizenship</strong></td><td><p>I am a Canadian citizen. </p><p></p><p>However I have worked for US companies through agencies like Toptal, EvolveSquads, and others.</p></td><td></td></tr><tr><td><strong>Security Clearance</strong></td><td>Significantly higher than Enhanced Reliability. Talk to me for details. </td><td></td></tr><tr><td><strong>Insurance</strong></td><td>I have Professional Liability Insurance and I am a member of the APCC (Association of Professional Canadian Consultants).</td><td></td></tr><tr><td><strong>Certifications</strong></td><td>CKAD (Jan 2021)</td><td></td></tr></tbody></table>


# Marketing & BD

If you have a product or service you would like to present to Sentian Cloud Computing Inc, please consider the following:

* I accept LinkedIn connection requests only ***from people I have met in person, by phone, or by video*** (such as Zoom, Meets, etc), and only if it makes sense. I never blindly accept a connection request, or just based on a common interest.&#x20;
* **I do not buy products or services** except what is strictly necessary to run my business: office equipment, computers, software engineering / devops tools, communication tools, management tools, online presence tools. My clients and the open source community provide the rest. My business does not sell products or SaaS, only "work-as-a-service".&#x20;
* I cannot promote your product or service unless I truly believe it is a better solution. This could only happen if I use it significantly and regularly; I cannot rely on your marketing material or a demo. Moreover, I will point out its strengths ***and*** weaknesses, and contrast them to alternatives, some of which could be your competition... since ultimately I want my clients to get what is best for them, without any bias on my part.&#x20;

Link back to my [LinkedIn](https://linkedin.com/in/oliver-schoenborn).


# What is DevOps

This is a question I get asked regularly. It always leads to an interesting discussion, because it is not a rigorously defined term. What prompted this post is an answer that I posted on LinkedIn to for that exact question: I realized I should gather some thoughts on the topic here.&#x20;

DevOps is primarily about bringing Operations and Development closer to accelerate delivery and deployment of computer applications. This is achieved by:&#x20;

1. coding configuration of infrastructure, and of delivery and deployment processes;&#x20;
2. using a common set of tools in Ops and Dev;&#x20;
3. using runtime environments in Dev that are representative of those in Ops;
4. aiming for immutability, idempotency, repeatability, and replicability of infrastructure and runtime environments.

Some devops enthusiasts add to the above that DevOps must be achieved via one team that does both Dev and Ops. I don't agree that it is a must: for one thing, I have only seen this work in small, early stage development projects. Mostly because DevOps tends to make use of a rather large set of specialized tools, certainly more as the system gets larger, and at some point it becomes impossible for everyone on a team to stay comfortable with all the aspects I mentioned above. Inevitably, some specialization and focus occurs for each member of the team.&#x20;

What is critical is that the devops team communicate with -- and work closely with -- all the dev teams, and enable them as much as possible to manage their own development environments, and to help them make their application deployments as close as feasible to the production deployments. &#x20;

DevOps is not system and/or network administration or Operations: these roles have a very different focus, namely administration of machines (physical or virtual), networks, and client-facing application, respectively.&#x20;


# Learning Programming

Having a mentor and then eventually, being a mentor, is a great way to learn. Nothing is quite as powerful as having to answer someone's questions about a programming language (or any technology or "system", for that matter), to deepen one's understanding of that language. AND you're helping someone else!\
\
If this interests you, have a look at [https://exercism.io](https://exercism.io/). This pairs you with mentor, or you can yourself be a mentor. There are over 50 programming languages to choose from!

Oliver Schoenborn


# Learning Go

I posted recently about exercism.io as a great way to learn a programming language, or to deepen one's understanding of a language by being a mentor.

More specifically to learn Go, here are 3 other sites that I recommend:

* tour of go: <https://tour.golang.org>
* play with go: <https://play-with-go.dev>
* gophercises: <https://courses.calhoun.io>

Oliver Schoenborn


# Docker run: Mirror Host User

Docker makes it easy to share host network and filesystem, but it doesn't make it easy to share the host's user ID and group ID. This is very useful when using a container locally via `docker run`, when it needs to write files to a volume shared with the host. In that case it is useful for the files to have the same user ID and group ID as the user on host.

Currently the way I do this is as follows:

* In the `Dockerfile`, I create a user `runuser` and group `rungroup` via `useradd` and `groupadd` commands, and I set final `USER` so container defaults to that user:

```
RUN groupadd rungroup \
 && useradd -ms /bin/bash -g rungroup runuser
 
USER runuser

...more setup...

# ensure container runs as runuser
USER runuser
```

* In a small script, run the docker container in detached mode (`--detach`)

```
  docker run \
    ...
    --env USER \
    --name $CTNR_NAME \
    --rm \
    --detach \
    --tty \
    IMAGE_NAME
```

* Then docker exec `usermod` and `groupmod` to match the host's user ID and group ID:&#x20;

```
docker exec -it -u root $CTNR_NAME groupmod -g "$GROUP_ID" rungroup
docker exec -it -u root $CTNR_NAME usermod -u $UID runuser
```

* A final line in the script executes the desired command in the container, such as a shell. The command will run as `runuser:rungroup` but with the ID that match that of the host:&#x20;

```
# runs shell as last USER in Dockerfile
docker exec -it $CTNR_NAME id
```

* If you need a root shell in container, say to install more apps temporarily (because if you restart the container those apps will be gone -- which is a good thing, ensure clean slate for any new container), change user:&#x20;

```
docker exec -it -u root $CTNR_NAME /bin/bash
```

This is quite tricky and required a fair bit of time to figure out.

I've seen another solution of mounting /etc/passwd and /etc/group but this exposes way more info in the container than necessary (which is just one line of each file). So for me this is not a solution.

In any case the above enables multiple simultaneous shells, each one running as either `root` or `runuser`, in latter case the UID and GROUP\_ID will match that of host user who started the container, and all shells can be exited without terminating the container.&#x20;

One caveat is that there will still be files owned by the original user ID that got created in the `Dockerfile`. Eg if the `useradd` command in `Dockerfile` created user `runuser` with ID 2000, and then in `Dockerfile` other commands are run as that user that creates files, the files will have ownership by user ID 2000. The `docker exec` that is run later changes the `runuser` ID to something else, but this does not change the ownership of any files already created. Therefore, you may need to chown those files via an additional docker exec. Eg&#x20;

```
docker exec -it -u root $CTNR_NAME \
  chown runuser /var/run/docker.sock
```

In the small script I additionally have a check to determine if the container is already running, in that case it skips the docker run, and also to easy choose between `runuser` and `root`:&#x20;

```
#!/usr/bin/env bash

run_as_root=false
if [[ ${1:-} == '--su' ]]; then
  shift
  echo "Will run as root"
  run_as_root=true
fi

CTNR_NAME=something

if [[ -z $( docker ps -qf name=$CTNR_NAME ) ]]; then
  echo "Starting new container $CTNR_NAME"
  docker run \
    ...

  # match host user ID and group ID
  docker exec -it -u root $CTNR_NAME groupmod -g "$GROUP_ID" rungroup
  docker exec -it -u root $CTNR_NAME usermod -u $UID runuser
  
  # some files need to be re-owned by runuser
  docker exec -it -u root $CTNR_NAME chown runuser /var/run/docker.sock
  
else
  echo "Container $CTNR_NAME is already running"
fi

echo "Shelling into $CTNR_NAME container"
if [[ $run_as_root == true ]]; then
  docker exec -it --user root    $CTNR_NAME /bin/bash
else
  docker exec -it --user runuser $CTNR_NAME /bin/bash
fi
```


# GitOps: thoughts, challenges, etc

GitOps is a great idea at its core: describe the desired state of  infrastructure & deployments through files and scripts in git, and have a program running somewhere that monitors changes to these files; when they change, this program takes action.&#x20;

&#x20;the actual infrastructure / deployment state vs this desired state and takes the desired state.&#x20;

Are there pros and cons to gitops, gray areas, gotchas, challenges?&#x20;

I'm not going to provide a \*definitive\* answer to this question because the landscape is constantly changing. So I'm just going to provide some thoughts that represent my own journey through gitops. I may update this post as my journey evolves, or I might write a new post, I'll see!

Aspects to consider:&#x20;

* If you want to be able to replicate an environment (for troubleshooting, testing etc), you need \*everything\* to be captured in git \*in some fashion\* (TBD below).&#x20;
  * This means not just the kubernetes manifests for your micro-service; it also includes the tools used to activate the target state (such as kubectl, helm, spinnaker).&#x20;
  * Here \*in some fashion\* should be interpreted liberally: it could mean eg that you have a Dockerfile that builds a docker image with the specific versions of the tools, AND your gitops uses that container to reach the desired state, AND it also records what docker tag of that container was used in the metadata of the cluster.&#x20;
* Gitops is often discussed in the context of kubernetes because the controller model of kubernetes is a natural fit for gitops. HOWEVER, in reality there are many deployment-related resources that live outside of a kubernetes cluster:&#x20;
  * databases and other forms of storage (blobs like S3, caching services like AWS redis)
  * message queues like SQS
  * lambdas / serverless functions&#x20;
  * security groups / firewall rules
  * networking eg a micro service being extended might need a new VPC endpoint
  * etc.&#x20;
* The non-cluster infra described in previous point must be versioned too, and similarly anything not micro-service code.&#x20;
  * docker files
  * test files
  * helm charts
  * values files
  * kustomize files
  * terraform files
  * k8s manifests
  * jenkins files and pipeline definitions in general
  * pulumi / Python scripts
  * cloudformation templates
  * OPA / policy files
  * etc
* A lot of git commits are typically needed on a feature breanch before a PR is needed. Yet developers need to run the compute associated with these commits in a representative "QA" or "sandbox" environment, without waiting to get to PR stage. So there has to be a mechanism for developers to deploy new code into live sandbox without a PR.  Yet, it is not safe to give them carte blanche either, because kubernetes provides access to huge compute resources which could get activated even inadvertently (by both junior and experienced, for a variety of reasons -- to err is human!).&#x20;
* When you have multiple developers working on features, they need a way to know what failed in their deployment.&#x20;
  * This means that the gitops system must capture the logs of its deployments so that they are still accessible after a failed service deployment gets rolled back, so the next developer can deploy into a working environment.&#x20;
  * What if the developer needs to troubleshoot? Clearly they can't deploy into the integration environment of the CI/CD. They need to be able to create an environment on demand and deploy into it.&#x20;
  * How are they going to know the context of that environment, ie the versions of all the other pods that were running at the moment that the test failed in the gitops/ci/cd deployment? The gitops must capture, upon a failure of deployment / test, the versions of all other services running. Moreover, the state of the databases must be captured, even possibly the state of caches. This is getting hard!
* Running something from the command line should be allowed, but there should be a sentinel that alerts if drift is not removed after a certain amount of time (eg an hour).&#x20;
* There does not seem to be a software that does all of the above
* The software available to tackle some of the above are rather big and involve a significant amount of ramp-up time with understanding how to map your requirements to their capabilities to their declarative language. Not to mention, if something does not work, these tools involve an additional layer to troubleshoot.&#x20;
* Intuit apparently uses 3 repos per micro-service! Although one of the 3 seems to be spring-boot specific.&#x20;
* The infrastructure related to gitops must be re-instated if the cluster gets lost / damaged!&#x20;
  * This means that the gitops setup must also be captured in git. And if you have 100 clusters, do you really want each one to have its own gitops controller and dashboard, and config in a git repo?&#x20;
  * It seems that it would be more effective to have one cluster dedicated to gitops, which can therefore have different life cycle requirements and constraints than the deployment environments, and would present one dashboard. In that case, the gitops agent would get a user/group in k8s RBAC.&#x20;
* Keeping the infra code separate from the micro-service git repo has pros but also cons: the build artifact may require changes to the infra code that will be done asynchronously hence for a while, the gitops will attempt to deploy a docker image that cannot work and will have to roll back, this is wasteful and adds audit noise
* The use of git as an event "bus" seems to complicate the gitops architecture. Eg I have built a simple gitops system that uses one git repo per micro-service, contains the helm chart and the values and sops-encoded secrets for the different environments (qa, staging, prod).&#x20;
  * In that case, there is no need for gitops tools to write to git repos, which in turn enables one repo rather than 2; the helm chart / values can be updated at same time as the micro-service, the PR is a great place to see what the developer is proposing, everything is in one place, close to the code. The test code and data is in the git repo; so should the ci/cd code and data!
  * Some might argue that this is dangerous, and that you want to control what happens at the infra level. I argue that the spirit of devops is to narraw the gap between dev and ops; now gitops via tools like flux and argocd seem to be widening that gap by requiring separate repos that can have PRs for infra etc. Seems very onerous. &#x20;
* Ironically, pull requests are not native to git. They are a controlled merge, based  on a service provided by a third-party such as github, bitbucket, gitlab etc.&#x20;
* <https://github.com/open-gitops/documents/blob/v0.1.0/PRINCIPLES.md>

About Weaveworks paper "[Automating Kubernetes with GitOps](https://go.weave.works/rs/249-YDT-025/images/Whitepaper_AutomatingKuberneteswithGitOps.pdf)":

* Stronger security guarantees:&#x20;
  * most orgs that I know do not use signed commits.&#x20;
  * git is not completely immutable; like anything it can be hacked, controls must be setup (typically manually for each repo), and history can be rewritten (eg rebase, reset, push force, etc). Since git becomes the entry point to deployment into some clusters, it is just a matter of time before these aspects get exploited.&#x20;
  * "separation of responsibility between packaging software and releasing it to a production environment embodies the security principle of least privilege": this is true, but can also increase the divide between dev and ops. Only weakest link needs to be exploited since gitops creates a chain that can end in prod just by checking in code, no need to know anything about the deployment environments (passwords, commands to run, registries, etc); in fact, the gitops minimizes (literally) the amount of knowledge an attacker needs in order to change the system.&#x20;
* Reduced mean time to recovery: assumes many of the things I described earlier are in git too!
* What you need for gitops:
  * declarative description of entire system; as describe above, this also include third-party dependencies including the gitops setup itself!
  * ability to auto apply approved changes: yes but re "you don’t need specific cluster credentials to make a change to your system. With GitOps, there is a segregated environment that the state definition lives outside of. This allows your team to separate what they actually do from how they are going to do it." I don't totally buy that approach, it requires extra git repos, it interferes with dev agility, it widens gap between dev and ops, it can cause undeployable artifact
* Seems to focus on PR as the way to validate changes but in reality, PR usually near end of feature branch, yet changes may need to be validated before the branch is PR'd as described earlier
* CICD pipeline:&#x20;
  * integration used to refer to the notion of bringing together many components of a system for testing;&#x20;
  * integ tests in a micro-services env require a live env due to the distributed architecture (at very least, a subset of a container's dependencies to be running in other containers in same network)
  * therefore the integ tests should actually be after the orchestrator, ie from git commit you build the docker image and run unit tests in it (the image is the unit of test), push it to image repo, deploy it via orchestrator, run integration tests, if deployment or tests fail the orchestrator must rollback, whereas if all succeeds there may be further actions such as push image to a "release" repo and/or do a deployment to a staging and/or prod (if tests in staging work).&#x20;
  * database migrations may be required, but in automated pipeline that must also be automatically triggered; and although they should always be backwards compatible, errors happens and someone is bound to release a service that runs a db migrations that breaks backwards compat and then what happens, how do you deal with this? especially when there is more than one branch being modified in git repo.&#x20;
  * security problems: CI need not deploy; you can separate ie trigger on new docker image, helm chart, etc. Also, the API creds used by CI tooling will be needed by the CD running inside the cluster, so whereas there is one way to gain access to CI, there are many to gain access to CD (by cracking any of the micro services running in the same cluster as the CD).&#x20;
  * cluster goes down: rebuild all images? that's silly. Eg if you use jenkins to build and spinnaker to deploy, all you need to know is the commit hash that was deployed for each service, and re-run the deployment action manually. I build CD so that all deployment actions can be done in one go via a version manifest file. I believe this statement may be a shortcoming that stems from gitops doing both the building and the deployment. Eg it should be possible to tell a good gitops system to deploy the latest master commit of git repo A and B and C since the code is all there, there should be no need to rebuild anything.&#x20;
  * gitops deployment pipeline "automates complex error prone tasks like having to manually update YAML manifests": helm charts and kustomize files are all based on yaml, so anytime the configuration of a service changes (eg, a new feature requires a new env var to exist), there is a chance the YAML of the chart or kustomize will need modification and possibly refactoring (eg once you reach a point that there are many env vars that could be defined more simply via a loop, the helm chart template will have to be edited significantly, or the kustomize will have to be split heavily, and typically these changes will have to be done to many different charts -- that should be the case if you have naming and infra code conventions)
  * security wise the weakest link is that gitops as described in that doc needs to give write access to the git repos, I really dislike that approach, which seems to stem from the desire to separate dev from ops. With the technique i describe above, there is no need for auto write to git repos, thus improving security even further. Further, prod cluster is public facing and that is the cluster that has gitops with write access to git!!! I bet my money that this will cause much pain in the next few years as this gets exploited.&#x20;
  * the separation between CI and CD-via-gitops as separation between dev and deployments is actually incorrect: integration comes from joining together pieces into a system, and this requires deployment; the whole notion that CI and CD are two separate things is in fact on shaky grounds in the world of containerized services.&#x20;
  * the gitops workflow described does not match my experience; as described earlier, a typical developer will need to deploy their service in a live environment where they can "integrate" it with real (or at least mock or fixture) DB and/or upstream/downstream services; and they will need to do that without a PR on their branch.&#x20;
  * pages 12-14 are a marketing pitch for Weaveworks Enterprise K8S Platform (nothing wrong with that, just nothing in those pages that is relevant to this journey).&#x20;


# View approved PRs in github

The PR dashboard in github does not show which ones were approved or are still waiting for approval. Two useful searches in github to get around this (if you work on several repos and branches):

Find all your own PRs (you are author) across all git repos, that have been approved:&#x20;

`org:YourOrgName is:pr is:open author:@me review:approved`

Find all your own PRs that have not yet been approved:&#x20;

`org:YourOrgName is:pr is:open author:@me -review:approved`


# Configuration for Terraform

In several of my projects, a variables.tf file with a set of basic variables is not adequate for configuration. An example is where a stack consists of multiple independent building blocks that each have their own configuration parameters. In that case, each module needs its own configuration file. Moreover, it can be useful to use a tree structure to represent configuration elements related to the same concepts. And finally, many of these have default values: simple default values, like if a stack does not specify a bucket name, the module tf code will generate one; and more complex types, eg the module may allow one element to be a map of arbitrary key names, but the values all have the same structure, in which case each map value has its own set of defaults.&#x20;

Ideally a configuration schema would have default values or indicate that a property is required, would have type info, support some value constraints (like via a regexp), and support defaults for basic data structures like strings, numbers, lists and maps. It might look something like this:&#x20;

```yaml
server: 
  size: string = "t2.nano"
  ssh_key_name: string = "server_key"
  
buckets:
  _map_key_bucket_id: 
    name: string = required
    policy_filename: string = null (auto-generated)
    retention_days: number = 7
    
cidr_blocks: 
  _list_:
    ip: regexp (\d{1,3}\.){3}\d{1,3} = 0.0.0.0
    bits: number = 24
    description: text = null
```

This is not currently available so I use a scheme where each re-usable terraform module has json/yaml defaults, and these can be overridden in a separate "overrides" file when the module is imported into a stack. The merged config is output by terraform.&#x20;

This works really well. EXCEPT that&#x20;

1. The IDE has no knowledge of this configuraiton system so while editing module code that references config tree, my IDE does not give me intellisense on the configuration tree. This becomes more of an issue as the system grows.&#x20;
2. I can't easily inject dynamically generated defaults into the defaults. Eg it would be nice if I could show that when config parameter "instance\_name" is null, then this means "use the name computed by the module's tf code". But terraform is not designed to generate object trees like that, so this type of processing is painful in HCL.&#x20;
3. Merging is too basic in terraform, and the excellent third-party module I use is not parsed properly by infracost.&#x20;

A custom provider could solve the #2 and 3, but not #1. So in my next iteration I will create a wrapper that generates the necessary tf files containing variable definitions with defaults, specific to the set of modules used for a stack. I might base this on one of the open source ones that exist like terragrunt, terraspace, terramate, cdktf, or atmos, I have to investigate.


# Storing Sensitive Data in Git

### Git

Git is easily used as the one source of truth when it comes to configuration data, esp. in the world of infrastructure as code. A good way to include sensitive config info in a git repo is via the [sops CLI](https://github.com/mozilla/sops).&#x20;

Basically the workflow is like this (using AWS KMS for encryption; sops also supports GCP KMS, Azure Key Vault, age, and PGP):

1. someone creates an AWS KMS key and adds AWS users to it; this only needs to be done once
2. any one of those users encrypts the sensitive config file using sops CLI and the created AWS KMS key, and pushes it to git
3. any one of those users who has checked out the git repo uses sops to decrypt the file (this will fail if user does not have permission to use one of the keys mentioned in the encrypted file).&#x20;

Sops puts the key reference in the encoded file, so decoding it is trivial, and safe: only users listed in the key (when the key was created) can use the key, so only they can decode it.&#x20;

The sops CLI purposely encourages to decrypt to memory or stdout; you have to use extra arguments to encrypt to file. So it is very easy to see the content of the file, decrypted, without writing to disk, thus decreasing likelihood of leakage.

Several keys can be used in a key ring to encrypt, for redundancy (ie if an AWS region becomes unavailable, you can still decrypt if you used keys from several regions); sops will automatically try each key until one is found that works.

Any number of files can use the same KMS key.

### Terraform

Sops is easy to use with terraform thanks to the [sops terraform provider](https://registry.terraform.io/providers/carlpett/sops/latest/docs) which I have used quite a bit and is hugely popular.

The workflow of the preceding section has a different step 3:&#x20;

3\. someone configures the terraform code to use the sops provider and a data source that loads the file, pushes to git

4\. some checks out the git repo

5\. When terraform plan or apply is run, the sops provider automatically uses the KMS key referenced in the encrypted file to decode the file in memory; the tf code can use this info

This will fail if the user running terraform is not in the list of users of the key in AWS.

The decoded config file is never written to disk.


# Resizing an Ubuntu Guest VM Disk

Just today I needed to expand a disk again on a couple of VirtualBox VMs running Ubuntu 20.04. The docs on Ubuntu say there is a `Disks` app that provides a GUI. But even after resizing the virtual hard disk in the VM media manager, and following those steps, `Disks` did not resize it, nor did it show any error.

&#x20;![](https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2FYXS5TON6TLmecGlZSedE%2Fimage.png?alt=media\&token=c7ea2098-d189-4b79-82d3-8eea288c98e8)![](https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2FHd6tdFz6kRUc158b9pe8%2Fimage.png?alt=media\&token=a7140002-b179-4786-af26-62250c1ded06)

Specifically, after clicking the red `Resize` button, nothing happened.&#x20;

A couple years ago I wrote a post on dev.to, "[Expand disk in virtualbox Ubuntu 18 guest](https://dev.to/schollii/expand-disk-in-virtualbox-ubuntu-18-guest-24eo)". The steps were basically: shutdown VM, resize hard disk in VM manager, boot VM, resize extended partition using gparted GUI, then 3 more complicated command line commands involving pvresize (physical volume resize), lvresize (logical volume resize), and resize2fs.&#x20;

But with gparted 1.0, it is possible to do everything after booting in the GUI, which is a lot simpler than my original post on dev.to. Here are the steps:&#x20;

### Resize virtual disk in VirtualBox

1. Power down the VM
2. Go to Virtual Media Manager (VMM)
3. Select the disk that corresponds to your VM and resize the disk file (this is possible only on powered down VM)
4. Start VM

### Resize partition in guest using gparted:

1. If you don't have gparted 1.0, install `gparted`: `sudo apt install gparted`. As of June 30, 2023 this installed 1.0.0.
2. Backup any data, in case something goes wrong. I just cloned the disk in virtualbox media manager.&#x20;
3. Start `gparted`
4. Select the extended partition (`/dev/sda2` in my case, see snapshot). This extended partition must be resized first, as it is the "physical partition" and contains your actual root logical partition (`/dev/sda5` in my case). Right-click it and select Resize...

<figure><img src="https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2Fmi1LKwL9umAHQaRdiPFv%2Fimage.png?alt=media&amp;token=2bf2740d-7ee1-4cde-a473-4ed78eb4ad58" alt="" width="375"><figcaption></figcaption></figure>

5. You can enter any number larger than the total size of the hard disk, gparted will automatically compute the correct max value. Click OK.&#x20;
   * After you OK this, you will see a pending operation in the bottom section, and the unallocated portion will now be child of the extended partition.&#x20;
   * You will also notice how the dashed box, which represents that parent partition, now includes the unallocated portion.&#x20;

<figure><img src="https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2FcNQOSJjUeEYIAzmoJTB5%2Fimage.png?alt=media&amp;token=76b29164-fd76-4d1d-a87c-432a86180e9c" alt="" width="375"><figcaption></figcaption></figure>

6. Now select the actual root partition to resize.

<figure><img src="https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2F4aah0bm1bfBwc1kXzRyB%2Fimage.png?alt=media&amp;token=0e17661d-f46d-4708-9e62-5e3a66e747c2" alt="" width="375"><figcaption></figcaption></figure>

7. Right-click and select Resize, this time you will have the option to increase logical volume size. Enter a large value, let gparted compute the max. Click OK. You will now see a second pending operation, and the graph reflects the new (unapplied) size.&#x20;

<figure><img src="https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2FCqt6YJprlfX3R1IUkpam%2Fimage.png?alt=media&amp;token=73f9208b-1aa6-474b-a648-4db285799fa8" alt="" width="375"><figcaption></figcaption></figure>

8. When you are ready, click the green checkmark to apply all pending operations. Success:

<figure><img src="https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2FxxNvQalJcnN49y851b67%2Fimage.png?alt=media&amp;token=879e919c-8c18-42ba-ab87-e81d0bac2654" alt="" width="368"><figcaption></figcaption></figure>

Indeed, `df` now shows new size of root partition (`/dev/sda5` in my case), hurray!

<figure><img src="https://1899019720-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MNs5Wv9Z8JMY1K3-8GZ%2Fuploads%2Fppdr18vX7TTHyPAJJKku%2Fimage.png?alt=media&amp;token=b7bf93a1-5211-4462-819b-e93276b0f3e9" alt=""><figcaption></figcaption></figure>


