Compare commits

...
16 Commits
Author SHA1 Message Date
Alexander KukushkinandGitHub 381a5b80d2 Release 1.5.4 (#931)
* Bump version
* Update release notes
* Make it possible to configure registration of Service in Consul via env variables
2019-01-15 12:14:19 +01:00
Alexander KukushkinandGitHub 71dae6a905 Optionally consider node not healthy if it is not on the latest timeline (#892)
The latest timeline is calculated from the `/history` key in DCS. In case there is no such key or it contains some garbage we consider the node healthy.
Closes https://github.com/zalando/patroni/issues/890
2019-01-15 11:16:30 +01:00
Alexander KukushkinandGitHub cf34fb3934 Relax requirements on superuser credentials (#930)
libpq allows opening connection without explicitly specifying neither username nor password. Depending on situation it would rely either on `pgpass` file or trust authentication method in pg_hba.conf.

Since pg_rewind is also using libpq, it could work the same way.

Fixes https://github.com/zalando/patroni/issues/928
2019-01-15 11:15:35 +01:00
Alexander KukushkinandGitHub e080ded44b Make logging configurable via YAML file (#927)
It allows changing logging settings in runtime by updating config and doing reload or sending `SIGHUP` to the Patroni process.
Important! Environment configuration names related to logging were renamed and documentation accordingly updated. For compatibility reasons Patroni still accepts `PATRONI_LOGLEVEL` and `PATRONI_FORMAT`, but some other variables related to logging, which were introduced only
recently (between releases), will stop working. I think it is ok, since we didn't release the new version yet and therefore it is very unlikely that somebody is using them except authors of corresponding PRs.

Example of log section in the config file:
```yaml
log:
  dir: /where/to/write/patroni/logs  # if not specified, write logs to stderr
  file_size: 50000000  # 50MB
  file_num: 10  # keep history of 10 files
  dateformat: '%Y-%m-%d %H:%M:%S'
  loggers:  # increase log verbosity for etcd.client and urllib3
    etcd.client: DEBUG
    urllib3: DEBUG
```
2019-01-15 08:42:13 +01:00
jouirandAlexander Kukushkin dec3656f6e Redirect HTTPServer exceptions to logger (#900) (#925)
By default they were written to stdout
2019-01-15 08:37:06 +01:00
Alexander KukushkinandGitHub 1e2d89fa58 Apply 5 second backoff when loading global config up on start (#922)
It doesn't make much sense to hammer DCS when we just starting up.
Fixes https://github.com/zalando/patroni/issues/919
2019-01-14 14:55:56 +01:00
Alexander KukushkinandGitHub 994863c18d Refactor wait_for_user_backends_to_close method (#917)
1. Log only debug level messages on any kind of error
2. Update regexp for matching postgres aux processes to make it compatible with postgres 11

Fixes https://github.com/zalando/patroni/issues/914
2019-01-14 14:55:45 +01:00
Alexander KukushkinandGitHub 3fce982909 Set archive_mode to off during the custom bootstrap (#911)
We want to avoid archiving WALs and history files until the cluster is fully functional. It should really help if the custom bootstrap involves pg_upgrade.
2019-01-14 14:55:23 +01:00
Lucas CapistrantandAlexander Kukushkin d306092cbc Explicitly secure rw perms for recovery.conf at creation time (#910)
We don't want anybody except patroni/postgres user reading this file, it contains replication user and password.
2019-01-14 14:22:04 +01:00
Dmitry DolgovandAlexander Kukushkin 11f7ceb521 Do not check types of standby_cluster configuration (#924)
Simply allow valid keys
2019-01-14 14:16:15 +01:00
bradnicholsonandAlexander Kukushkin 05a13839aa Update replica_bootstrap.rst (#915)
Add some docs about replication slots for standby clusters
2019-01-14 12:57:21 +01:00
Étienne MandAlexander Kukushkin 04ac199fc8 Single quotes are mandatory around each host in PATRONI_ETCD_HOSTS (#926)
Otherwise YAML parser fails
2019-01-14 11:56:15 +01:00
anikin-aaandAlexander Kukushkin 0be8a9527b possibility to set logdatefmt via env. (#904)
The value could be configured via `PATRONI_LOGDATEFMT` environment variable. The default value is `%Y-%m-%d %H:%M:%S`
2019-01-14 08:46:38 +01:00
Pavel KirillovandAlexander Kukushkin 929ff08bfd Service deregister timeout must be in Go time format (#893) 2018-12-21 15:42:15 +01:00
Cody CoonsandAlexander Kukushkin 7bc8d0aac9 Removed stderr pipe to stdout on pg_ctl process (#896)
Inheriting stderr from the main Patroni process allows all Postgres logs to be seen along with all patroni logs. This is very useful in a container environment as Patroni and Postgres logs may be consumed using standard tools (docker logs, kubectl, etc).

In addition to that, this change fixes a bug with Patroni not being able to catch postmaster pid when postgres writing some warnings into stderr
2018-12-21 15:41:34 +01:00
Lucas CapistrantandAlexander Kukushkin f3da6de129 Add ability to configure app logs to be written to a file (#903)
It gives users the option to send Patroni application logs to a File instead of Standard Out. There are three environment variables that can be set to enable and configure file logging.
1. `PATRONI_FILE_LOG_DIR`: Path to a directory that is writeable by the executing user. Having this variable set is what activates file logging.
2. `PATRONI_FILE_LOG_NUM`: This is a rolling file logger. This variable dictates how many log files are retained.
3. `PATRONI_FILE_LOG_SIZE`: This variable dictates the size at which the logs will roll.

If `PATRONI_FILE_LOG_DIR` is not set than Patroni will log to stderr (default behavior does not change)

Closes https://github.com/zalando/patroni/issues/902
2018-12-21 15:38:29 +01:00
24 changed files with 371 additions and 83 deletions
+3
View File
@@ -51,3 +51,6 @@ scm-source.json
docs/build/
docs/source/_static/
docs/source/_templates/
# Pycharm IDE
.idea/
+11 -4
View File
@@ -11,8 +11,13 @@ Global/Universal
- **PATRONI\_NAME**: name of the node where the current instance of Patroni is running. Must be unique for the cluster.
- **PATRONI\_NAMESPACE**: path within the configuration store where Patroni will keep information about the cluster. Default value: "/service"
- **PATRONI\_SCOPE**: cluster name
- **PATRONI\_LOGLEVEL**: sets the general logging level (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
- **PATRONI\_REQUESTS_LOGLEVEL**: sets the logging level for all HTTP requests e.g. Kubernetes API calls (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
- **PATRONI\_LOG\_LEVEL**: sets the general logging level. Default value is **INFO** (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
- **PATRONI\_LOG\_FORMAT**: sets the log formatting string. Default value is **%(asctime)s %(levelname)s: %(message)s** (see `the LogRecord attributes <https://docs.python.org/3.6/library/logging.html#logrecord-attributes>`_)
- **PATRONI\_LOG\_DATEFORMAT**: sets the datetime formatting string. (see the `formatTime() documentation <https://docs.python.org/3.6/library/logging.html#logging.Formatter.formatTime>`_)
- **PATRONI\_LOG\_DIR**: Directory to write application logs to. The directory must exist and be writable by the user executing Patroni. If you set this env variable, the application will retain 4 25MB logs by default. You can tune those retention values with `PATRONI_LOG_FILE_NUM` and `PATRONI_LOG_FILE_SIZE` (see below).
- **PATRONI\_LOG\_FILE\_NUM**: The number of application logs to retain.
- **PATRONI\_LOG\_FILE\_SIZE**: Size of patroni.log file (in bytes) that triggers a log rolling.
- **PATRONI\_LOG\_LOGGERS**: Redefine logging level per python module. Example ``PATRONI_LOG_LOGGERS="{patroni.postmaster: WARNING, urllib3: DEBUG}"``
Bootstrap configuration
-----------------------
@@ -36,11 +41,13 @@ Consul
- **PATRONI\_CONSUL\_KEY**: (optional) File with the client key. Can be empty if the key is part of certificate.
- **PATRONI\_CONSUL\_DC**: (optional) Datacenter to communicate with. By default the datacenter of the host is used.
- **PATRONI\_CONSUL\_CHECKS**: (optional) list of Consul health checks used for the session. If not specified Consul will use "serfHealth" in additional to the TTL based check created by Patroni. Additional checks, in particular the "serfHealth", may cause the leader lock to expire faster than in `ttl` seconds when the leader instance becomes unavailable.
- **PATRONI\_CONSUL\_REGISTER\_SERVICE**: (optional) whether or not to register a service with the name defined by the scope parameter and the tag master, replica or standby-leader depending on the node's role. Defaults to **false**
- **PATRONI\_CONSUL\_SERVICE\_CHECK\_INTERVAL**: (optional) how often to perform health check against registered url
Etcd
----
- **PATRONI\_ETCD\_HOST**: the host:port for the etcd endpoint.
- **PATRONI\_ETCD\_HOSTS**: list of etcd endpoints in format host1:port1,host2:port2,etc...
- **PATRONI\_ETCD\_HOSTS**: list of etcd endpoints in format 'host1:port1','host2:port2',etc...
- **PATRONI\_ETCD\_URL**: url for the etcd, in format: http(s)://(username:password@)host:port
- **PATRONI\_ETCD\_PROXY**: proxy url for the etcd. If you are connecting to the etcd using proxy, use this parameter instead of **PATRONI\_ETCD\_URL**
- **PATRONI\_ETCD\_SRV**: Domain to search the SRV record(s) for cluster autodiscovery.
@@ -79,7 +86,7 @@ PostgreSQL
- **PATRONI\_SUPERUSER\_PASSWORD**: password for the superuser, set during initialization (initdb).
REST API
--------
--------
- **PATRONI\_RESTAPI\_CONNECT\_ADDRESS**: IP address and port to access the REST API.
- **PATRONI\_RESTAPI\_LISTEN**: IP address and port that Patroni will listen to, to provide health-check information for HAProxy.
- **PATRONI\_RESTAPI\_USERNAME**: Basic-auth username to protect unsafe REST API endpoints.
+12
View File
@@ -10,6 +10,18 @@ Global/Universal
- **namespace**: path within the configuration store where Patroni will keep information about the cluster. Default value: "/service"
- **scope**: cluster name
Log
---
- **level**: sets the general logging level. Default value is **INFO** (see `the docs for Python logging <https://docs.python.org/3.6/library/logging.html#levels>`_)
- **format**: sets the log formatting string. Default value is **%(asctime)s %(levelname)s: %(message)s** (see `the LogRecord attributes <https://docs.python.org/3.6/library/logging.html#logrecord-attributes>`_)
- **dateformat**: sets the datetime formatting string. (see the `formatTime() documentation <https://docs.python.org/3.6/library/logging.html#logging.Formatter.formatTime>`_)
- **dir**: Directory to write application logs to. The directory must exist and be writable by the user executing Patroni. If you set this value, the application will retain 4 25MB logs by default. You can tune those retention values with `file_num` and `file_size` (see below).
- **file\_num**: The number of application logs to retain.
- **file\_size**: Size of patroni.log file (in bytes) that triggers a log rolling.
- **loggers**: This section allows redefining logging level per python module
- **patroni.postmaster: WARNING**
- **urllib3: DEBUG**
Bootstrap configuration
-----------------------
- **dcs**: This section will be written into `/<namespace>/<scope>/config` of a given configuration store after initializing of new cluster. This is the global configuration for the cluster. If you want to change some parameters for all cluster nodes - just do it in DCS (or via Patroni API) and all nodes will apply this configuration.
+1
View File
@@ -76,6 +76,7 @@ Also, the following Patroni configuration options can be changed only dynamicall
- loop_wait: 10
- retry_timeouts: 10
- maximum_lag_on_failover: 1048576
- check_timeline: false
- postgresql.use_slots: true
Upon changing these options, Patroni will read the relevant section of the configuration stored in DCS and change its
+59
View File
@@ -3,6 +3,65 @@
Release notes
=============
Version 1.5.4
-------------
This version implements flexible logging and fixes a number of bugs.
**New features**
- Improvements in logging infrastructure (Alexander Kukushkin, Lucas Capistrant, Alexander Anikin)
Logging configuration could be configured not only from environment variables but also from Patroni config file. It makes it possible to change logging configuration in runtime by updating config and doing reload or sending SIGHUP to the Patroni process. By default Patroni writes logs to stderr, but now it becomes possible to write logs directly into the file and rotate when it reaches a certain size. In addition to that added support of custom dateformat and the possibility to fine-tune log level for each python module.
- Make it possible to take into account the current timeline during leader elections (Alexander Kukushkin)
It could happen that the node is considering itself as a healthiest one although it is currently not on the latest known timeline. In some cases we want to avoid promoting of such node, which could be achieved by setting `check_timeline` parameter to `true` (default behavior remains unchanged).
- Relaxed requirements on superuser credentials
Libpq allows opening connections without explicitly specifying neither username nor password. Depending on situation it relies either on pgpass file or trust authentication method in pg_hba.conf. Since pg_rewind is also using libpq, it will work the same way.
- Implemented possibility to configure Consul Service registration and check interval via environment variables (Alexander Kukushkin)
Registration of service in Consul was added in the 1.5.0, but so far it was only possible to turn it on via patroni.yaml.
**Stability Improvements**
- Set archive_mode to off during the custom bootstrap (Alexander Kukushkin)
We want to avoid archiving wals and history files until the cluster is fully functional. It really helps if the custom bootstrap involves pg_upgrade.
- Apply five seconds backoff when loading global config on start (Alexander Kukushkin)
It helps to avoid hammering DCS when Patroni just starting up.
- Reduce amount of error messages generated on shutdown (Alexander Kukushkin)
They were harmless but rather annoying and sometimes scary.
- Explicitly secure rw perms for recovery.conf at creation time (Lucas)
We don't want anybody except patroni/postgres user reading this file, because it contains replication user and password.
- Redirect HTTPServer exceptions to logger (Julien Riou)
By default, such exceptions were logged on standard output messing with regular logs.
**Bug fixes**
- Removed stderr pipe to stdout on pg_ctl process (Cody Coons)
Inheriting stderr from the main Patroni process allows all Postgres logs to be seen along with all patroni logs. This is very useful in a container environment as Patroni and Postgres logs may be consumed using standard tools (docker logs, kubectl, etc). In addition to that, this change fixes a bug with Patroni not being able to catch postmaster pid when postgres writing some warnings into stderr.
- Set Consul service check deregister timeout in Go time format (Pavel Kirillov)
Without explicitly mentioned time unit registration was failing.
- Relax checks of standby_cluster cluster configuration (Dmitry Dolgov, Alexander Kukushkin)
It was accepting only strings as valid values and therefore it was not possible to specify the port as integer and create_replica_methods as a list.
Version 1.5.3
-------------
+2
View File
@@ -195,3 +195,5 @@ in a patroni configuration:
Note, that these options will be applied only once during cluster bootstrap,
and the only way to change them afterwards is through DCS.
If you use replication slots on the standby cluster, you must also create the corresponding replication slot on the primary cluster. It will not be done automatically by the standby cluster implementation. You can use Patroni's permenant replication slots feature on the primary cluster to maintain a replication slot with the same name as ``primary_slot_name``, or its default value if ``primary_slot_name`` is not provided.
+2
View File
@@ -13,6 +13,8 @@ In asynchronous mode the cluster is allowed to lose some committed transactions
The amount of transactions that can be lost is controlled via ``maximum_lag_on_failover`` parameter. Because the primary transaction log position is not sampled in real time, in reality the amount of lost data on failover is worst case bounded by ``maximum_lag_on_failover`` bytes of transaction log plus the amount that is written in the last ``ttl`` seconds (``loop_wait``/2 seconds in the average case). However typical steady state replication delay is well under a second.
By default, when running leader elections, Patroni does not take into account the current timeline of replicas, what in some cases could be undesirable behavior. You can prevent the node not having the same timeline as a former master become the new leader by changing the value of ``check_timeline`` parameter to ``true``.
PostgreSQL synchronous replication
----------------------------------
+1
View File
@@ -372,6 +372,7 @@ class ConsulController(AbstractDcsController):
def __init__(self, context):
super(ConsulController, self).__init__(context)
os.environ['PATRONI_CONSUL_HOST'] = 'localhost:8500'
os.environ['PATRONI_CONSUL_REGISTER_SERVICE'] = 'on'
self._client = consul.Consul()
self._config_file = None
+6 -6
View File
@@ -14,6 +14,7 @@ class Patroni(object):
from patroni.config import Config
from patroni.dcs import get_dcs
from patroni.ha import Ha
from patroni.log import PatroniLogger
from patroni.postgresql import Postgresql
from patroni.version import __version__
from patroni.watchdog import Watchdog
@@ -21,7 +22,9 @@ class Patroni(object):
self.setup_signal_handlers()
self.version = __version__
self.logger = PatroniLogger()
self.config = Config()
self.logger.reload_config(self.config.get('log', {}))
self.dcs = get_dcs(self.config)
self.watchdog = Watchdog(self.config)
self.load_dynamic_configuration()
@@ -49,6 +52,7 @@ class Patroni(object):
break
except DCSError:
logger.warning('Can not get cluster from dcs')
time.sleep(5)
def get_tags(self):
return {tag: value for tag, value in self.config.get('tags', {}).items()
@@ -65,6 +69,7 @@ class Patroni(object):
def reload_config(self):
try:
self.tags = self.get_tags()
self.logger.reload_config(self.config.get('log', {}))
self.dcs.reload_config(self.config)
self.watchdog.reload_config(self.config)
self.api.reload_config(self.config['restapi'])
@@ -138,12 +143,6 @@ class Patroni(object):
def patroni_main():
logformat = os.environ.get('PATRONI_LOGFORMAT', '%(asctime)s %(levelname)s: %(message)s')
loglevel = os.environ.get('PATRONI_LOGLEVEL', 'INFO')
requests_loglevel = os.environ.get('PATRONI_REQUESTS_LOGLEVEL', 'WARNING')
logging.basicConfig(format=logformat, level=loglevel)
logging.getLogger('requests').setLevel(requests_loglevel)
patroni = Patroni()
try:
patroni.run()
@@ -151,6 +150,7 @@ def patroni_main():
pass
finally:
patroni.shutdown()
logging.shutdown()
def pg_ctl_start(args):
+7
View File
@@ -3,6 +3,7 @@ import json
import logging
import psycopg2
import time
import traceback
import dateutil.parser
import datetime
import os
@@ -547,3 +548,9 @@ class RestApiServer(ThreadingMixIn, HTTPServer, Thread):
and self.__initialize(config):
self.start()
self.__set_config_parameters(config)
@staticmethod
def handle_error(request, client_address):
address, port = client_address
logger.warning('Exception happened during processing of request from {}:{}'.format(address, port))
logger.warning(traceback.format_exc())
+36 -18
View File
@@ -2,7 +2,6 @@ import json
import logging
import os
import shutil
import six
import sys
import tempfile
import yaml
@@ -44,6 +43,7 @@ class Config(object):
__DEFAULT_CONFIG = {
'ttl': 30, 'loop_wait': 10, 'retry_timeout': 10,
'maximum_lag_on_failover': 1048576,
'check_timeline': False,
'master_start_timeout': 300,
'synchronous_mode': False,
'synchronous_mode_strict': False,
@@ -195,12 +195,9 @@ class Config(object):
elif name not in ('connect_address', 'listen', 'data_dir', 'pgpass', 'authentication'):
config['postgresql'][name] = deepcopy(value)
elif name == 'standby_cluster':
allowed_keys = self.__DEFAULT_CONFIG['standby_cluster'].keys()
expected = {
k: v for k, v in (value or {}).items()
if (k in allowed_keys and isinstance(v, six.string_types))
}
config['standby_cluster'].update(expected)
for name, value in (value or {}).items():
if name in self.__DEFAULT_CONFIG['standby_cluster']:
config['standby_cluster'][name] = deepcopy(value)
elif name in config: # only variables present in __DEFAULT_CONFIG allowed to be overriden from DCS
if name in ('synchronous_mode', 'synchronous_mode_strict'):
config[name] = value
@@ -220,6 +217,15 @@ class Config(object):
if value:
ret[param] = value
def _fix_log_env(name, oldname):
value = _popenv(oldname)
name = Config.PATRONI_ENV_PREFIX + 'LOG_' + name.upper()
if value and name not in os.environ:
os.environ[name] = value
for name, oldname in (('level', 'loglevel'), ('format', 'logformat'), ('dateformat', 'log_datefmt')):
_fix_log_env(name, oldname)
def _set_section_values(section, params):
for param in params:
value = _popenv(section + '_' + param)
@@ -228,6 +234,22 @@ class Config(object):
_set_section_values('restapi', ['listen', 'connect_address', 'certfile', 'keyfile'])
_set_section_values('postgresql', ['listen', 'connect_address', 'data_dir', 'pgpass', 'bin_dir'])
_set_section_values('log', ['level', 'format', 'dateformat', 'dir', 'file_size', 'file_num', 'loggers'])
def _parse_dict(value):
if not value.strip().startswith('{'):
value = '{{{0}}}'.format(value)
try:
return yaml.safe_load(value)
except Exception:
logger.exception('Exception when parsing dict %s', value)
return None
value = ret.get('log', {}).pop('loggers', None)
if value:
value = _parse_dict(value)
if value:
ret['log']['loggers'] = value
def _get_auth(name):
ret = {}
@@ -266,22 +288,18 @@ class Config(object):
name, suffix = (param[8:].split('_', 1) + [''])[:2]
if name and suffix:
# PATRONI_(ETCD|CONSUL|ZOOKEEPER|EXHIBITOR|...)_(HOSTS?|PORT|..)
if suffix in ('HOST', 'HOSTS', 'PORT', 'SRV', 'URL', 'PROXY', 'CACERT', 'CERT',
'KEY', 'VERIFY', 'TOKEN', 'CHECKS', 'DC', 'NAMESPACE', 'CONTEXT',
'USE_ENDPOINTS', 'SCOPE_LABEL', 'ROLE_LABEL', 'POD_IP', 'PORTS', 'LABELS'):
if suffix in ('HOST', 'HOSTS', 'PORT', 'SRV', 'URL', 'PROXY', 'CACERT', 'CERT', 'KEY', 'VERIFY',
'TOKEN', 'CHECKS', 'DC', 'REGISTER_SERVICE', 'SERVICE_CHECK_INTERVAL', 'NAMESPACE',
'CONTEXT', 'USE_ENDPOINTS', 'SCOPE_LABEL', 'ROLE_LABEL', 'POD_IP', 'PORTS', 'LABELS'):
value = os.environ.pop(param)
if suffix == 'PORT':
value = value and parse_int(value)
elif suffix in ('HOSTS', 'PORTS', 'CHECKS'):
value = value and _parse_list(value)
elif suffix == 'LABELS':
if not value.strip().startswith('{'):
value = '{{{0}}}'.format(value)
try:
value = yaml.safe_load(value)
except Exception:
logger.exception('Exception when parsing dict %s', value)
value = None
value = _parse_dict(value)
elif suffix == 'REGISTER_SERVICE':
value = parse_bool(value)
if value:
ret[name.lower()][suffix.lower()] = value
# PATRONI_<username>_PASSWORD=<password>, PATRONI_<username>_OPTIONS=<option1,option2,...>
@@ -340,7 +358,7 @@ class Config(object):
'scope',
'retry_timeout',
'synchronous_mode',
'maximum_lag_on_failover'
'synchronous_mode_strict',
)
pg_config.update({p: config[p] for p in updated_fields if p in config})
+22 -2
View File
@@ -368,7 +368,7 @@ class SyncState(namedtuple('SyncState', 'index,leader,sync_standby')):
return name is not None and name in (self.leader, self.sync_standby)
class TimelineHistory(namedtuple('TimelineHistory', 'index,lines')):
class TimelineHistory(namedtuple('TimelineHistory', 'index,value,lines')):
"""Object representing timeline history file"""
@staticmethod
@@ -384,7 +384,7 @@ class TimelineHistory(namedtuple('TimelineHistory', 'index,lines')):
lines = None
if not isinstance(lines, list):
lines = []
return TimelineHistory(index, lines)
return TimelineHistory(index, value, lines)
class Cluster(namedtuple('Cluster', 'initialize,config,leader,last_leader_operation,members,failover,sync,history')):
@@ -484,6 +484,26 @@ class Cluster(namedtuple('Cluster', 'initialize,config,leader,last_leader_operat
slots = self.get_replication_slots(name, 'master').values()
return any(v for v in slots if v.get("type") == "logical")
@property
def timeline(self):
"""
>>> Cluster(0, 0, 0, 0, 0, 0, 0, 0).timeline
0
>>> Cluster(0, 0, 0, 0, 0, 0, 0, TimelineHistory.from_node(1, '[]')).timeline
1
>>> Cluster(0, 0, 0, 0, 0, 0, 0, TimelineHistory.from_node(1, '[["a"]]')).timeline
0
"""
if self.history:
if self.history.lines:
try:
return int(self.history.lines[-1][0]) + 1
except Exception:
logger.error('Failed to parse cluster history from DCS: %s', self.history.lines)
elif self.history.value == '[]':
return 1
return 0
@six.add_metaclass(abc.ABCMeta)
class AbstractDCS(object):
+2 -1
View File
@@ -389,7 +389,8 @@ class Consul(AbstractDCS):
api_parts = urlparse(data['api_url'])
api_parts = api_parts._replace(path='/{0}'.format(role))
conn_parts = urlparse(data['conn_url'])
check = base.Check.http(api_parts.geturl(), self._service_check_interval, deregister=self._client.http.ttl * 10)
check = base.Check.http(api_parts.geturl(), self._service_check_interval,
deregister='{0}s'.format(self._client.http.ttl * 10))
params = {
'service_id': '{0}/{1}'.format(self._scope, self._name),
'address': conn_parts.hostname,
+24 -5
View File
@@ -20,25 +20,28 @@ from threading import RLock
logger = logging.getLogger(__name__)
class _MemberStatus(namedtuple('_MemberStatus', 'member,reachable,in_recovery,wal_position,tags,watchdog_failed')):
class _MemberStatus(namedtuple('_MemberStatus', ['member', 'reachable', 'in_recovery', 'timeline',
'wal_position', 'tags', 'watchdog_failed'])):
"""Node status distilled from API response:
member - dcs.Member object of the node
reachable - `!False` if the node is not reachable or is not responding with correct JSON
in_recovery - `!True` if pg_is_in_recovery() == true
wal_position - value of `replayed_location` or `location` from JSON, dependin on its role.
timeline - timeline value from JSON
wal_position - maximum value of `replayed_location` or `received_location` from JSON
tags - dictionary with values of different tags (i.e. nofailover)
watchdog_failed - indicates that watchdog is required by configuration but not available or failed
"""
@classmethod
def from_api_response(cls, member, json):
is_master = json['role'] == 'master'
timeline = json.get('timeline', 0)
wal = not is_master and max(json['xlog'].get('received_location', 0), json['xlog'].get('replayed_location', 0))
return cls(member, True, not is_master, wal, json.get('tags', {}), json.get('watchdog_failed', False))
return cls(member, True, not is_master, timeline, wal, json.get('tags', {}), json.get('watchdog_failed', False))
@classmethod
def unknown(cls, member):
return cls(member, False, None, 0, {}, False)
return cls(member, False, None, 0, 0, {}, False)
def failover_limitation(self):
"""Returns reason why this node can't promote or None if everything is ok."""
@@ -92,6 +95,9 @@ class Ha(object):
def is_paused(self):
return self.check_mode('pause')
def check_timeline(self):
return self.check_mode('check_timeline')
def get_standby_cluster_config(self):
if self.cluster and self.cluster.config and self.cluster.config.modify_index:
config = self.cluster.config.data
@@ -543,15 +549,23 @@ class Ha(object):
:returns True when node is lagging
"""
lag = (self.cluster.last_leader_operation or 0) - wal_position
return lag > self.state_handler.config.get('maximum_lag_on_failover', 0)
return lag > self.patroni.config.get('maximum_lag_on_failover', 0)
def _is_healthiest_node(self, members, check_replication_lag=True):
"""This method tries to determine whether I am healthy enough to became a new leader candidate or not."""
_, my_wal_position = self.state_handler.timeline_wal_position()
if check_replication_lag and self.is_lagging(my_wal_position):
logger.info('My wal position exceeds maximum replication lag')
return False # Too far behind last reported wal position on master
if not self.is_standby_cluster() and self.check_timeline():
cluster_timeline = self.cluster.timeline
my_timeline = self.state_handler.replica_cached_timeline(cluster_timeline)
if my_timeline < cluster_timeline:
logger.info('My timeline %s is behind last known cluster timeline %s', my_timeline, cluster_timeline)
return False
# Prepare list of nodes to run check against
members = [m for m in members if m.name != self.state_handler.name and not m.nofailover and m.api_url]
@@ -562,11 +576,13 @@ class Ha(object):
logger.warning('Master (%s) is still alive', st.member.name)
return False
if my_wal_position < st.wal_position:
logger.info('Wal position of %s is ahead of my wal position', st.member.name)
return False
return True
def is_failover_possible(self, members):
ret = False
cluster_timeline = self.cluster.timeline
members = [m for m in members if m.name != self.state_handler.name and not m.nofailover and m.api_url]
if members:
for st in self.fetch_nodes_statuses(members):
@@ -575,6 +591,9 @@ class Ha(object):
logger.info('Member %s is %s', st.member.name, not_allowed_reason)
elif self.is_lagging(st.wal_position):
logger.info('Member %s exceeds maximum replication lag', st.member.name)
elif self.check_timeline() and (not st.timeline or st.timeline < cluster_timeline):
logger.info('Timeline %s of member %s is behind the cluster timeline %s',
st.timeline, st.member.name, cluster_timeline)
else:
ret = True
else:
+66
View File
@@ -0,0 +1,66 @@
import logging
import os
from copy import deepcopy
from logging.handlers import RotatingFileHandler
from patroni.utils import deep_compare
class PatroniLogger(object):
DEFAULT_LEVEL = 'INFO'
DEFAULT_FORMAT = '%(asctime)s %(levelname)s: %(message)s'
def __init__(self):
self.root_logger = logging.getLogger()
self.config = None
self.handler = None
self.reload_config({'level': 'DEBUG'})
def update_loggers(self):
loggers = deepcopy(self.config.get('loggers') or {})
for name, logger in self.root_logger.manager.loggerDict.items():
if not isinstance(logger, logging.PlaceHolder):
level = loggers.pop(name, logging.NOTSET)
logger.setLevel(level)
for name, level in loggers.items():
logger = self.root_logger.manager.getLogger(name)
logger.setLevel(level)
def reload_config(self, config):
if self.config is None or not deep_compare(self.config, config):
self.root_logger.setLevel(config.get('level', PatroniLogger.DEFAULT_LEVEL))
add_handler = None
if 'dir' in config:
if not isinstance(self.handler, RotatingFileHandler):
add_handler = RotatingFileHandler(os.path.join(config['dir'], __name__))
handler = add_handler or self.handler
handler.maxBytes = int(config.get('file_size', 25000000))
handler.backupCount = int(config.get('file_num', 4))
else:
if self.handler is None or isinstance(self.handler, RotatingFileHandler):
add_handler = logging.StreamHandler()
handler = add_handler or self.handler
oldlogformat = (self.config or {}).get('format', PatroniLogger.DEFAULT_FORMAT)
logformat = config.get('format', PatroniLogger.DEFAULT_FORMAT)
olddateformat = (self.config or {}).get('dateformat') or None
dateformat = config.get('dateformat') or None # Convert empty string to `None`
if oldlogformat != logformat or olddateformat != dateformat or add_handler:
handler.setFormatter(logging.Formatter(logformat, dateformat))
if add_handler:
self.root_logger.addHandler(add_handler)
if self.handler is not None:
self.root_logger.removeHandler(self.handler)
self.handler.close()
self.handler = add_handler
self.config = config.copy()
self.update_loggers()
+19 -16
View File
@@ -5,6 +5,7 @@ import re
import shlex
import shutil
import socket
import stat
import subprocess
import tempfile
import time
@@ -391,7 +392,7 @@ class Postgresql(object):
we have either wal_log_hints or checksums turned on
"""
# low-hanging fruit: check if pg_rewind configuration is there
if not (self.config.get('use_pg_rewind') and all(self._superuser.get(n) for n in ('username', 'password'))):
if not self.config.get('use_pg_rewind'):
return False
cmd = [self._pgcommand('pg_rewind'), '--help']
@@ -1110,8 +1111,12 @@ class Postgresql(object):
f.write(self._CONFIG_WARNING_HEADER)
f.write("include '{0}'\n\n".format(self.config.get('custom_conf') or self._postgresql_base_conf_name))
for name, value in sorted((configuration or self._server_parameters).items()):
if not self._running_custom_bootstrap or name != 'hba_file':
if not self._running_custom_bootstrap or name not in ('hba_file', 'archive_mode'):
f.write("{0} = '{1}'\n".format(name, value))
# we want to set archive_mode to 'off' during the custom bootstrap
# in order to avoid premature archiving of wals and history files
if self._running_custom_bootstrap:
f.write("archive_mode = 'off'\n")
# when we are doing custom bootstrap we assume that we don't know superuser password
# and in order to be able to change it, we are opening trust access from a certain address
# therefore we need to make sure that hba_file is not overriden
@@ -1154,8 +1159,10 @@ class Postgresql(object):
with open(self._pg_hba_conf, 'w') as f:
f.write(self._CONFIG_WARNING_HEADER)
for address, t in addresses.items():
f.write('{0}\t{1}\t{2}\t{3}\ttrust\n'.format(t, 'all',
self._superuser.get('username') or 'all', address))
f.write((
'{0}\treplication\t{1}\t{3}\ttrust\n'
'{0}\tall\t{2}\t{3}\ttrust\n'
).format(t, self._replication['username'], self._superuser.get('username') or 'all', address))
elif not self._server_parameters.get('hba_file') and self.config.get('pg_hba'):
with open(self._pg_hba_conf, 'w') as f:
f.write(self._CONFIG_WARNING_HEADER)
@@ -1186,6 +1193,7 @@ class Postgresql(object):
def write_recovery_conf(self, recovery_params):
with open(self._recovery_conf, 'w') as f:
os.chmod(self._recovery_conf, stat.S_IWRITE | stat.S_IREAD)
for name, value in recovery_params.items():
f.write("{0} = '{1}'\n".format(name, value))
@@ -1665,7 +1673,8 @@ $$""".format(name, ' '.join(options)), name, password, password)
def post_bootstrap(self, config, task):
try:
self.create_or_update_role(self._superuser['username'], self._superuser['password'], ['SUPERUSER'])
if 'username' in self._superuser and 'password' in self._superuser:
self.create_or_update_role(self._superuser['username'], self._superuser['password'], ['SUPERUSER'])
task.complete(self.run_bootstrap_post_init(config))
if task.result:
@@ -1684,17 +1693,11 @@ $$""".format(name, ' '.join(options)), name, password, password)
os.unlink(self._pg_hba_conf)
self.restore_configuration_files()
self._write_postgresql_conf()
if self._server_parameters.get('hba_file') and \
self._server_parameters['hba_file'] != self._pg_hba_conf:
self.restart()
else:
self._replace_pg_hba()
if self.pending_restart:
self.restart()
else:
self.reload()
time.sleep(1) # give a time to postgres to "reload" configuration files
self.close_connection() # close connection to reconnect with a new password
self._replace_pg_hba()
# at this point there should be no recovery.conf
if os.path.isfile(self._recovery_conf) or os.path.islink(self._recovery_conf):
os.unlink(self._recovery_conf)
self.restart()
except Exception:
logger.exception('post_bootstrap')
task.complete(False)
+23 -20
View File
@@ -102,28 +102,31 @@ class PostmasterProcess(psutil.Process):
return None
def wait_for_user_backends_to_close(self):
# These regexps are cross checked against versions PostgreSQL 9.1 .. 9.6
aux_proc_re = re.compile("(?:postgres:)( .*:)? (?:""(?:startup|logger|checkpointer|writer|wal writer|"
"autovacuum launcher|autovacuum worker|stats collector|wal receiver|archiver|"
"wal sender) process|bgworker: )")
# These regexps are cross checked against versions PostgreSQL 9.1 .. 11
aux_proc_re = re.compile("(?:postgres:)( .*:)? (?:(?:archiver|startup|autovacuum launcher|autovacuum worker|"
"checkpointer|logger|stats collector|wal receiver|wal writer|writer)(?: process )?|"
"walreceiver|wal sender process|walsender|walwriter|background writer|"
"logical replication launcher|logical replication worker for|bgworker:) ")
try:
user_backends = []
user_backends_cmdlines = []
for child in self.children():
try:
cmdline = child.cmdline()[0]
if not aux_proc_re.match(cmdline):
user_backends.append(child)
user_backends_cmdlines.append(cmdline)
except psutil.NoSuchProcess:
pass
if user_backends:
logger.debug('Waiting for user backends %s to close', ', '.join(user_backends_cmdlines))
psutil.wait_procs(user_backends)
logger.debug("Backends closed")
children = self.children()
except psutil.Error:
logger.exception('wait_for_user_backends_to_close')
return logger.debug('Failed to get list of postmaster children')
user_backends = []
user_backends_cmdlines = []
for child in children:
try:
cmdline = child.cmdline()[0]
if not aux_proc_re.match(cmdline):
user_backends.append(child)
user_backends_cmdlines.append(cmdline)
except psutil.NoSuchProcess:
pass
if user_backends:
logger.debug('Waiting for user backends %s to close', ', '.join(user_backends_cmdlines))
psutil.wait_procs(user_backends)
logger.debug("Backends closed")
@staticmethod
def start(pgcommand, data_dir, conf, options):
@@ -157,7 +160,7 @@ class PostmasterProcess(psutil.Process):
cmdline = [pgcommand, '-D', data_dir, '--config-file={}'.format(conf)] + options
logger.debug("Starting postgres: %s", " ".join(cmdline))
proc = call_self(['pg_ctl_start'] + cmdline, close_fds=(os.name != 'nt'),
stdout=subprocess.PIPE, stderr=subprocess.STDOUT, env=env)
stdout=subprocess.PIPE, env=env)
pid = int(proc.stdout.readline().strip())
proc.wait()
logger.info('postmaster pid=%s', pid)
+1 -1
View File
@@ -1 +1 @@
__version__ = '1.5.3'
__version__ = '1.5.4'
+4 -1
View File
@@ -73,7 +73,7 @@ class MockHa(object):
@staticmethod
def fetch_nodes_statuses(members):
return [_MemberStatus(None, True, None, None, {}, False)]
return [_MemberStatus(None, True, None, 0, None, {}, False)]
@staticmethod
def schedule_future_restart(data):
@@ -401,3 +401,6 @@ class TestRestApiServer(unittest.TestCase):
self.assertRaises(ValueError, srv.reload_config, bad_config)
self.assertRaises(ValueError, srv.reload_config, {})
srv.reload_config({'listen': '127.0.0.2:8008'})
def test_handle_error(self):
self.assertIsNone(MockRestApiServer.handle_error(None, ('127.0.0.1', 55555)))
+16 -1
View File
@@ -1,6 +1,6 @@
import os
import unittest
import sys
import unittest
from mock import MagicMock, Mock, patch
from patroni.config import Config
@@ -30,6 +30,8 @@ class TestConfig(unittest.TestCase):
'PATRONI_NAME': 'postgres0',
'PATRONI_NAMESPACE': '/patroni/',
'PATRONI_SCOPE': 'batman2',
'PATRONI_LOGLEVEL': 'ERROR',
'PATRONI_LOG_LOGGERS': 'patroni.postmaster: WARNING, urllib3: DEBUG',
'PATRONI_RESTAPI_USERNAME': 'username',
'PATRONI_RESTAPI_PASSWORD': 'password',
'PATRONI_RESTAPI_LISTEN': '0.0.0.0:8008',
@@ -49,6 +51,7 @@ class TestConfig(unittest.TestCase):
'PATRONI_ETCD_CERT': '/cert',
'PATRONI_ETCD_KEY': '/key',
'PATRONI_CONSUL_HOST': '127.0.0.1:8500',
'PATRONI_CONSUL_REGISTER_SERVICE': 'on',
'PATRONI_KUBERNETES_LABELS': 'a:b:c',
'PATRONI_KUBERNETES_SCOPE_LABEL': 'a',
'PATRONI_KUBERNETES_PORTS': '[{"name": "postgresql"}]',
@@ -84,3 +87,15 @@ class TestConfig(unittest.TestCase):
self.config.save_cache()
with patch('os.fdopen', MagicMock()):
self.config.save_cache()
def test_standby_cluster_parameters(self):
dynamic_configuration = {
'standby_cluster': {
'create_replica_methods': ['wal_e', 'basebackup'],
'host': 'localhost',
'port': 5432
}
}
self.config.set_dynamic_configuration(dynamic_configuration)
for name, value in dynamic_configuration['standby_cluster'].items():
self.assertEqual(self.config['standby_cluster'][name], value)
+13 -5
View File
@@ -28,8 +28,10 @@ def false(*args, **kwargs):
def get_cluster(initialize, leader, members, failover, sync, cluster_config=None):
history = TimelineHistory(1, [(1, 67197376, 'no recovery target specified', datetime.datetime.now().isoformat())])
cluster_config = cluster_config or ClusterConfig(1, {1: 2}, 1)
t = datetime.datetime.now().isoformat()
history = TimelineHistory(1, '[[1,67197376,"no recovery target specified","' + t + '"]]',
[(1, 67197376, 'no recovery target specified', t)])
cluster_config = cluster_config or ClusterConfig(1, {'check_timeline': True}, 1)
return Cluster(initialize, cluster_config, leader, 10, members, failover, sync, history)
@@ -72,12 +74,13 @@ def get_standby_cluster_initialized_with_only_leader(failover=None, sync=None):
)
def get_node_status(reachable=True, in_recovery=True, wal_position=10, nofailover=False, watchdog_failed=False):
def get_node_status(reachable=True, in_recovery=True, timeline=2,
wal_position=10, nofailover=False, watchdog_failed=False):
def fetch_node_status(e):
tags = {}
if nofailover:
tags['nofailover'] = True
return _MemberStatus(e, reachable, in_recovery, wal_position, tags, watchdog_failed)
return _MemberStatus(e, reachable, in_recovery, timeline, wal_position, tags, watchdog_failed)
return fetch_node_status
@@ -115,6 +118,7 @@ zookeeper:
sys.argv = sys.argv[:1]
self.config = Config()
self.config.set_dynamic_configuration({'maximum_lag_on_failover': 5})
self.postgresql = p
self.dcs = d
self.api = Mock()
@@ -147,6 +151,7 @@ def run_async(self, func, args=()):
@patch.object(Postgresql, 'query', Mock())
@patch.object(Postgresql, 'checkpoint', Mock())
@patch.object(Postgresql, 'cancellable_subprocess_call', Mock(return_value=0))
@patch.object(Postgresql, '_get_local_timeline_lsn_from_replication_connection', Mock(return_value=[2, 10]))
@patch.object(etcd.Client, 'write', etcd_write)
@patch.object(etcd.Client, 'read', etcd_read)
@patch.object(etcd.Client, 'delete', Mock(side_effect=etcd.EtcdException))
@@ -166,7 +171,6 @@ class TestHa(unittest.TestCase):
mock_machines.__get__ = Mock(return_value=['http://remotehost:2379'])
self.p = Postgresql({'name': 'postgresql0', 'scope': 'dummy', 'listen': '127.0.0.1:5432',
'data_dir': 'data/postgresql0', 'retry_timeout': 10,
'maximum_lag_on_failover': 5,
'authentication': {'superuser': {'username': 'foo', 'password': 'bar'},
'replication': {'username': '', 'password': ''}},
'parameters': {'wal_level': 'hot_standby', 'max_replication_slots': 5, 'foo': 'bar',
@@ -461,6 +465,8 @@ class TestHa(unittest.TestCase):
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.fetch_node_status = get_node_status(watchdog_failed=True)
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.fetch_node_status = get_node_status(timeline=1)
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
self.ha.fetch_node_status = get_node_status(wal_position=1)
self.assertEqual(self.ha.run_cycle(), 'no action. i am the leader with the lock')
# manual failover from the previous leader to us won't happen if we hold the nofailover flag
@@ -572,6 +578,8 @@ class TestHa(unittest.TestCase):
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
with patch('patroni.postgresql.Postgresql.timeline_wal_position', return_value=(1, 1)):
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
with patch('patroni.postgresql.Postgresql.replica_cached_timeline', return_value=1):
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
self.ha.patroni.nofailover = True
self.assertFalse(self.ha._is_healthiest_node(self.ha.old_cluster.members))
self.ha.patroni.nofailover = False
+38
View File
@@ -0,0 +1,38 @@
import os
import sys
import unittest
import yaml
from mock import Mock, patch
from patroni.config import Config
from patroni.log import PatroniLogger
class TestPatroniLogger(unittest.TestCase):
@patch('logging.FileHandler._open', Mock())
def setUp(self):
self.config = {
'log': {
'dir': 'foo',
'file_size': 4096,
'file_num': 5,
'loggers': {
'foo.bar': 'INFO'
}
},
'restapi': {}, 'postgresql': {'data_dir': 'foo'}
}
sys.argv = ['patroni.py']
os.environ[Config.PATRONI_CONFIG_VARIABLE] = yaml.dump(self.config, default_flow_style=False)
self.logger = PatroniLogger()
config = Config()
self.logger.reload_config(config['log'])
def test_rotating_handler(self):
self.assertEqual(self.logger.handler.maxBytes, self.config['log']['file_size'])
self.assertEqual(self.logger.handler.backupCount, self.config['log']['file_num'])
def test_reload_config(self):
self.config['log'].pop('dir')
self.logger.reload_config(self.config['log'])
+1
View File
@@ -651,6 +651,7 @@ class TestPostgresql(unittest.TestCase):
@patch('time.sleep', Mock())
@patch('os.unlink', Mock())
@patch('os.path.isfile', Mock(return_value=True))
@patch.object(Postgresql, 'run_bootstrap_post_init', Mock(return_value=True))
@patch.object(Postgresql, '_custom_bootstrap', Mock(return_value=True))
@patch.object(Postgresql, 'start', Mock(return_value=True))
+2 -3
View File
@@ -67,7 +67,7 @@ class TestPostmasterProcess(unittest.TestCase):
@patch('psutil.wait_procs')
def test_wait_for_user_backends_to_close(self, mock_wait):
c1 = Mock()
c1.cmdline = Mock(return_value=["postgres: startup process"])
c1.cmdline = Mock(return_value=["postgres: startup process "])
c2 = Mock()
c2.cmdline = Mock(return_value=["postgres: postgres postgres [local] idle"])
c3 = Mock()
@@ -77,8 +77,7 @@ class TestPostmasterProcess(unittest.TestCase):
self.assertIsNone(proc.wait_for_user_backends_to_close())
mock_wait.assert_called_with([c2])
c3.cmdline = Mock(side_effect=psutil.AccessDenied(123))
with patch('psutil.Process.children', Mock(return_value=[c3])):
with patch('psutil.Process.children', Mock(side_effect=psutil.NoSuchProcess(123))):
proc = PostmasterProcess(123)
self.assertIsNone(proc.wait_for_user_backends_to_close())